The last few months in AI have moved so fast that it is hard to keep track of everything. In September, OpenAI announced that its agents had found a blowup solution to the 3D Navier-Stokes equations, a result tied to one of the Millennium Prize Problems. Then in early October, OpenAI posted more than 700 AI-generated math preprints in one go, which one physicist called a "slopocalypse". Every few weeks a new model comes out that feels like another step on the path towards recursive self-improvement, and the breakthroughs keep arriving faster than most of us can even read about them.

AI math milestones went from a year apart to two weeks apart Selected announcements, July 2024 to October 2026 Jul 2024 Jan 2025 Jul 2025 Jan 2026 Jul 2026 Jul 25, 2024 IMO silver (AlphaProof) Jul 21, 2025 IMO gold (Gemini Deep Think) Jan 4, 2026 First Erdős problem solved May 20, 2026 Unit distance conjecture disproved Sep 8, 2026 Navier-Stokes blowup claim Sep 21, 2026 100+ more open problems Oct 6, 2026 700+ AI math preprints
Sources are Semafor, Google DeepMind, Quanta Magazine, TechCrunch and Nature, linked in the references below.

So it is no surprise that there is a lot of doom around academia right now. When agents can search the literature, run experiments and write up full papers on their own, people naturally start asking whether there is any point in doing a PhD anymore, or in becoming a researcher at all, if a system can do most of the work faster than you can.

And yet some of the most exciting stories this year point the other way. Earlier this month Pavel Rabtsevich, a researcher working on his own without a lab or funding behind him, found a promising new planet candidate hiding in seven years of NASA's TESS telescope data. He used Claude Code with Opus 5.5 and Fable 5.1 to write more than a thousand scripts that downloaded the data and searched and vetted every signal, and NASA has since approved telescope time to check what he found. The planet still has to be confirmed, but what got him this far was not an institution or a big budget. It was a good taste for data and the curiosity to keep asking the next question while agents did the heavy lifting.

Lately a lot of juniors and early-career researchers have been reaching out to me with some version of the same question, which is how do you even do research in a world of automated AI agents and systems. I understand where it comes from. When I started doing research full time a year ago, AI tools were already really useful for coding and for the literature survey, but they were nowhere close to fully automated systems that could run a research project end to end. The real shift came after Anthropic released Claude Fable 5 in June this year. Since then the way I research, the way I choose questions and the timelines I work on have changed far more than I expected.

When I joined Lossfunk as a research engineer last year, my days were pretty simple. I would go into the lab, read papers and spend a lot of time discussing them with my peers, talking through what we found interesting, what seemed to be missing and where an idea could go next. Looking back, those conversations mattered as much as the reading itself, because talking things through with people who think differently from you is how half-formed thoughts turn into ideas worth exploring, and that is where a lot of my early thinking about research really took shape.

My view is that research hasn't become any less valuable, but the skills that matter have shifted. The answer is not to resist AI agents, and it is not to hand everything over to them either. It is to embrace them as collaborators, spend real time understanding what they give you and do research with them rather than against them. This post is my attempt to share what I have learned about that so far, mostly by getting things wrong along the way.

What changed this year

When I started, AI tools were mostly assistants that gave me a faster search or a quick summary of a paper I had already found. Today agents can run the whole loop on their own. They find and read the papers, write the code, run the experiments and draft the write-up.

That has changed how I spend my time. I collect far fewer papers and spend much more time judging what is actually worth trusting. The timelines have also compressed a lot, so an idea that felt novel at the start of a project might not be novel anymore by the time you finish it.

The biggest change is in which questions are worth picking. With a proper harness around them, AI systems can now automate a lot of the low-hanging fruit in research, like quick ablations and incremental follow-up experiments that an agent can grind through on its own. So the real work for a researcher has moved to finding an important research question with lots of possibilities in it. That means something open enough to explore in many directions and hard enough that it needs your judgment at every step.

The literature survey is where it hurts most

The stage that has changed the most is the literature survey. There is now far more content coming in than any researcher has the time to read, understand and actually absorb, and a good part of it is AI-generated research and slop papers that look polished but don't add much. The tricky part is that a weak paper takes just as long to read as a good one, and it is getting harder to tell them apart from the outside.

To put that in numbers, NeurIPS alone went from about 2,400 main track submissions in 2016 to more than 30,000 in 2026. And when the organizers checked this year's position paper track, 273 of the 969 submissions, about 28%, came back flagged as substantially AI-written.

NeurIPS submissions grew nearly 13x in a decade Main track submissions, 2,406 in 2016 to 30,709 in 2026 2016: 2,406 submissions 2016 2017: 3,240 submissions 2017 2018: 4,856 submissions 2018 2019: 6,743 submissions 2019 2020: 9,467 submissions 2020 2021: 9,122 submissions 2021 2022: 10,411 submissions 2022 2023: 12,343 submissions 2023 2024: 15,671 submissions 2024 2025: 21,575 submissions 2025 2026: 30,709 submissions 2026 2,406 30,709
NeurIPS main track statistics via OpenAccept and CSConfStats, 2016 to 2026.

The old approach of reading everything relevant and then finding the gap simply doesn't work when the pile of everything grows faster than you can read it. So I feel the skill now isn't about reading more, it is about deciding what deserves your attention and what you can safely skip.

How I decide what to read

I don't try to read everything anymore. Instead I go through a few stages, where each one takes a bit more time than the one before it.

  1. Almost every paper now comes with a research thread on X or a short post, so I start there and ask myself whether the problem actually sticks with me, and whether it is something unique and useful or I just got attracted to a catchy idea.
  2. If it does stick, I pass the paper through an LLM and ask for the abstract, the main takeaways and where the idea could go further, which gives me a good overview in a few minutes.
  3. If I still find it interesting after that, I sit down and do a proper deep read of the paper from start to end.
Most papers stop at the thread, only a few earn a deep read Read the thread Does the problem stick? LLM pass Abstract and takeaways Deep read Start to end, properly Skip it and move on sticks still interesting no no
How I filter papers, in three stages.

Most papers stop at the first stage, and that is completely fine, because it means the few that make it to the deep read actually get the time and focus they deserve.

Spend your time on the hard problems

With research getting automated and slop piling up everywhere, I feel focus has become the most valuable thing a researcher has. Agents have made it cheap to produce papers, but they haven't made it any cheaper to produce important ones, so it is really important to put your time into the tough, ambitious and useful problems rather than just getting more papers out.

To me, ambitious doesn't mean making a small tweak to an existing algorithm or approach. Take LLM harnesses, which are really popular right now. Building yet another harness or adding one more feature on top of an existing one is the easy move, and the harder and more interesting move is to look one level deeper at the problems that harnesses are trying to work around. There are a few of these that I keep coming back to.

  • How models can reason over long horizons, and whether they can build efficient abstractions or even world models while working through long tasks
  • How to solve the context compaction problem, particularly for small LLMs
  • Whether we can disentangle the knowledge stored in a model from the genuine reasoning it has actually learned

None of these has a clean answer yet, and some might stay a work in progress for a long time, but that is exactly what makes them worth your time. They give you real hypotheses to test, and you can use agents to go through related papers, work out possible ideas and run the experiments. If it works out, you have helped the community make progress on a problem that really matters, and if it doesn't, you have still learned a lot and built a deeper intuition by trying to solve something difficult and cool.

Taste comes from questions, not papers

Picking the right problem needs research taste, and there isn't really a shortcut to building it. It comes from reading papers, going through the relevant literature and slowly understanding why certain problems matter more than others. In my first couple of months, most of my time went into reading and discussing papers before I really understood the issues that are most important in AI, and I still think that time was one of the best investments I made.

Taste builds up slowly and in layers. You read the initial papers on a problem, then the follow-ups that push on them, and somewhere along the way you start running your own experiments to see what actually holds up. Each of those steps adds a little intuition about what works and what is worth chasing, and over time that intuition turns into research taste, which is really just a sense for which problems are the most important ones to attempt.

For me one such example was EsoLang-Bench. I kept wondering whether models that solve coding problems are actually generalizing or just reproducing what they have memorized, so I used esoteric programming languages that almost nobody writes code in as a way to tell the two apart.

The point isn't the paper that came out of it. It is learning to ask important research questions, building intuition around them and working on them wholeheartedly whether the research pans out or not, because even the experiments that fail end up deepening your understanding of the problem.

Work with agents, not against them

This year so much has changed, with AI agents now coming up with research questions to investigate, doing all the coding and testing and even writing out the paper. It is easy to react to this by either ignoring agents or fearing them, and I think both are a mistake. The researchers who will do the best work are the ones who learn to collaborate with these systems, treating them as a partner that is fast and tireless but still needs a human who understands the problem and can judge what comes back.

Say a paper on context compaction for small models makes it through all three of your reading stages. An agent can help you understand it much more deeply than a first read would, by questioning its assumptions and results along with you. It can help you come up with future work questions that build on top of it and turn your rough hunches into hypotheses you can actually test, and it can write the code for the training or inference experiments, which saves a lot of time.

Own everything the agent gives you

The catch is that you have to use agents responsibly, which means checking every output you get, going through the code, debugging it and making sure you actually understand what it is doing.

That work is how you develop intuition and taste in your own research, and it is what helps you in the long run. Without it you end up putting out work that you haven't really understood or done yourself, which can become a real problem for your credibility as a researcher. It also takes away the very thing that makes you better over time, which is the ability to identify the tough problems and come up with solutions for them.

Agents speed up the loop, but the question and the check stay yours Ask a hard question One you really care about Read and discuss Papers, threads, peers Form hypotheses Turn hunches into tests Agents build and run Code, training, inference Verify and debug Read every line they return Share in public Posts, threads, discussion new questions never hand these over hand this to agents, then check it
The research loop I follow, in six steps.

Share your work in public

Research doesn't end when the paper is out, and I feel sharing your work in public through blog posts, threads and discussions is just as essential as doing the work itself.

The same threads I use to filter papers are how other people will end up finding yours. Putting your work out there gets you feedback early, helps surface the flaws you might have missed and connects you with people who are working on the same problems. It also sharpens your own thinking, because if you can't explain an idea plainly to someone else, you probably don't understand it well enough yet. And in a feed full of slop, a clear explanation from the person who actually did the work really stands out. It doesn't have to start online either, since those lab discussions with my peers were really the first version of this, and sharing in public is just the same habit at a bigger scale.

The fun part is still yours

If I were starting again today, this is roughly what I would do.

  1. Spend the first few months reading deeply and understanding which problems in the field really matter.
  2. Filter papers hard, starting with the thread, then an LLM pass, and only doing a deep read for the ones that stick with you.
  3. Pick one hard question you care about and commit to it, even if it might not pan out.
  4. Use agents to question papers, test ideas and write your experiment code.
  5. Review, debug and understand everything they hand back to you.
  6. Share what you learn in public through blog posts, threads and discussions.

Agents can take over the summaries, the boilerplate code and the first draft, but don't hand over the part where you choose which question deserves your year and understand your answer well enough to stand behind it. The future of research is humans and AI working together, so embrace these systems, spend time with what they produce and keep yourself in the loop. That is the whole fun of research, and agents can help you get there faster as long as you don't use them to skip it.

References

  1. Semafor, "Google DeepMind AI system reaches milestone in global math contest", 25 July 2024 (AlphaProof, IMO silver).
  2. Google DeepMind, "Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad", 21 July 2025.
  3. Quanta Magazine, "Why the legendary Erdős problems are falling to AI", 3 August 2026.
  4. Quanta Magazine, "AI has solved one of math's $1 million Millennium Prize Problems", 8 September 2026.
  5. TechCrunch, "OpenAI forms math advisory group as its AI resolves more than 100 open problems", 21 September 2026.
  6. Nature, "OpenAI posts 700 maths preprints online", 7 October 2026.
  7. Tech Brew, "Anthropic's fabled release", June 2026 (Claude Fable 5).
  8. Pavel Rabtsevich, "AI-assisted TESS exoplanet candidate TIC 4206066", October 2026.
  9. HuggingNews, "Researcher finds 10 more exoplanets using Opus 5.5 AI", October 2026.
  10. NeurIPS Blog, "AI-generated papers in the NeurIPS 2026 position paper track", 2 June 2026.
  11. OpenAccept and CSConfStats, NeurIPS submission statistics, 2016 to 2026.