Some thoughts on theoretical research in 2026
Where are we headed, and why?
Harold Cohen’s “AARON Gijon”, 2007. AARON was one of “the earliest artificial intelligence (AI) program for artmaking”, according to the Whitney Museum of American Art; “Cohen understood his work with AARON to be a collaboration, and he devoted his life to exploring the potential of artificial intelligence to translate an artist’s knowledge and process into code.”
Disclosure: I have never worked on “core” TCS, only a bit on learning theory, and some related concepts in the algorithmic fairness literature. However, I do consider a fair amount of the research I performed to be theoretically oriented, and hence, I feel like I have a reasonable base of experience from which to draw the ideas in this post. I also recommend reading Timothy Gower’s recent post, and Terence Tao’s recent presentation on similar topics, both of which are surely more eloquent and well-thought through than this post.
Hsin-Yuan Huang recently led a very nice discussion at the UC Berkeley Simons Institute on the future of TCS research (see minute 24 onward for the relevant bit). The audience posed many insightful questions that articulate broader feelings that I have been experiencing over the past six months, and I think pretty much everyone agrees that these are urgent questions for the TCS community to wrestle with. I figured that I could share some constructive thoughts on some of these questions, if only for me to develop them further.
All discussions are generally predicated on the following nearly incontrovertible fact: AI systems are now solving research level TCS / theory problems “one-shot” simply by prompting them to do so. See, e.g., this recent set of results. This is simultaneously frightening and exhilarating to those who have spent years working on problems in TCS or related areas.
I have sampled a few questions and themes from the audience of the talk that I think are quite important.
What is the value of an academic paper if it can be produced one-shot by an LLM with a simple prompt?
Following that, given that our academic and cultural institutions around research are predicated on “papers” as units of currency, we will surely need to rebuild the way research is evaluated and performed. How should this be done?
What will TCS research even look like in the near future?
The Role of Academic Papers
These themes and questions certainly heightened my own uncertainty and anxiety about continuing in TCS, and contributed to the fact that I chose not to pursue a postdoc after my PhD. In particular, I believe that the depreciation of the “academic paper” as a unit of currency does not bode well for the near-term continuation of the current iteration of the academic institution, at least for Computer Science. In a few years we will have figured out what to do as a community, but the next (0, 3] years seem like they may be quite tumultuous.
As a side note, I think there is of course a distinction between a regular paper and a good paper. It is difficult to articulate what a regular and good paper are. However, I think most authors would be able to point to which of their works fall into each category.1
One could argue that number of regular papers has already ceased to be a useful proxy for the quality of an individual candidate, and that what really matters is number of good papers. People like saying this, but I don’t believe it is completely true for at least three different reasons.
If the number of regular papers increases exponentially (see below figure), it becomes harder for good papers to rise to the top of the stack, buried under an ever growing amount of noise.
Arxiv submissions by year: https://arxiv.org/stats/monthly_submissions
Every researcher can pinpoint what they would consider to be their “best” papers, i.e., those that they spent the most time and effort working on, or the results they are most proud of cracking. In my experience, however, this can often have little correlation with what work the community may most appreciate, at least for conventional metrics such as citations, or what work other people in the community think of when they see your name.
Finally, traditionally “successful” academics and new university hires aren’t — at least in TCS / ML — typically candidates with only K high quality papers for low K. They all usually have a large number of papers, with a couple standouts of high quality.
All these mean that you need to play the game of a high # of regular papers in addition to a few high quality good papers.
More broadly, academia has always had credentialism at its core. Previously, that included a multitude of factors such as advisor, academic institution, papers, awards, etc. Removing papers from this formula might, at least in the short term, push academia to revert to an even stronger reliance on traditional credentials, name-brand institutions, and status as academic micro-celebrities who can ensure their work gets read and shared widely due to their name and popularity.
As an example of this changing credentialism structure, ICLR (and probably other conferences) will implement restrictions on who can submit papers, effectively “locking out” new ML researchers who have never previously submitted papers.
Papers as a Unit of Work
I have never worked in “core” TCS; most of my experience comes from within the learning theory and algorithmic fairness sub-communities. There, many individuals worked on classical PAC learning and variants thereof, or different theoretical notions of algorithmic fairness.
One thing that eventually made me quite cynical about these kinds of papers in particular is that it often felt like a small group of researchers were applying the same bag of tricks to very similar problems, whose statements changed only in simply stated parameters. I very rarely felt that these papers exposed something fundamental about the world. Is the world a better place for the 1015th bandit variation paper?2
Nonetheless, if you are an author who constructed and checked the proofs or arguments for any of these theorems and got them past peer review, this functioned as a “stamp” on your intellectual credentials and CV. In theory, the very fact that you spent a non-trivial amount of time thinking about this problem verified that you at least had some level of intellectual capability.
A good fraction of this sort of research may simply be eliminated due to AI. In particular, many of these simple problem variations / re-parametrizations have proofs that certainly can be one-shotted by GPT 5.6, especially if they would have yielded due to a common trick from the bag.
Returning to Huang’s talk, one audience member — unfortunately I could not identify them — raised a very interesting point which sort of went unaddressed. Paraphrasing, they argued that all “theorems” can themselves be considered as “datapoints” of facts which are actually true about the world. According to the above, some of these datapoints may be more interesting than others. Prior to automation of the proofs of theorems, the number of datapoints we could generate as a community was bounded by the number of researchers actively working in related areas. However, it is conceivable to now imagine that a single TCS graduate student can, with the aid of GPT-X, produce hundreds or thousands of theorems over the course of their PhD. Each of these should obviously not be a single paper anymore.
So, what is the new unit of work for researchers? Or what should it be?
Ideas for New “Units of Work”
Around two to three months ago, I recommended that my graduating undergraduate students who were heading off to their PhD programs should become familiar with Lean and using AI tools like GPT for proving theorems. I think this is obvious now, and, if I were starting my PhD today, I might actually radically re-think what the main output of my PhD should be.
One short-term vision I have about the value of humans (at least for the next few months) is that we can function as guides for the AI, turning their powers and tools to shape the direction of a particular subfield. If one PhD student can indeed “prove” hundreds or thousands of theorems with the help of GPT (with little of their own intellectual contribution), then I think there are at least two remaining roles left.
Maybe the place that they apply their judgement is in understanding and guiding the advancements of the AI in order to generate new, more important, and more difficult conjectures to prove.
Providing expository mathematical writing, and unearth the “key” ideas of proofs. These “key” ideas seem to be the important bits that let us better understand truths in the world.
One example of the latter is provided by Thomas Bloom (maintainer of the Erdos problems website), when he talks about two “AI era” math researchers, Liam and Kevin.
There are issues with this “reframing” of the academic contribution of a PhD student or researcher. The primary issue is that institutions are slow to adapt to changes in the standard “you must publish K papers to be successful / graduate” requirement. Nonetheless, I strongly believe that early PhD students should not worry about fulfilling the minimum requirements necessary, and forge their own path and experiment here. I don’t think you have much to lose!3
Here is an example of what such a “radical” approach may look like. I will use the “multicalibration” literature as an example (see this great ICML tutorial by Ira, Natalie, and Aaron), since I am familiar with it, but the details of the field are not important for this description. Here’s what I would want to start with.
First, collect all the published, theoretical papers on multicalibration.
Spend some time formalizing all of these papers with AI, creating some sort of “Lean database” of all theorems in the literature. We may already find some issues or gaps in the literature just by doing this!
Now the fun part, use AI to help you think about and discover new connections and conjectures, weed out low hanging fruit (probably a lot of my own papers), and help surface interesting insights which may be useful to a broader audience.
Point 3 is where your clarity of thought, research taste, and expository technical writing ability will help you stand out. Why is a particular theorem interesting? What does it reveal about some fundamental tradeoff in the world? Why is the proof revealing? A researcher is then judged on their contributions in advancing the entire subfield. The bar can be much higher now! Did they spend years thinking about the entire subfield deeply? Are they the leading expert in this subfield? Is the subfield itself important to the world? This all seems a bit more “subjective” than just producing a large number of papers, and as such, our evaluation mechanisms may need to evolve.
Furthermore, if everything is formalized in a modular way within a Lean database, the database itself could then become a resource that everyone in the community works from and contributes to. What used to be an academic “paper” becomes some sort of blog post or released code artifact, and discussions can actually happen even more collaboratively online.
Increased Paper Quality Thresholds
I think the following thing is also likely to happen with respect to papers. Since they are so central to the academic institution, they are unlikely to disappear completely. Instead, the amount of “work” that should be contained in a single paper will increase. That is, a single or couple theorems will not be sufficient condition for acceptance. Instead, the paper will have to create “tangible” or “significant” impact on an entire subfield.
This may implicitly be enforced by what I see as inevitable upcoming paper submission quotas. For example, if each author can only submit five papers to ICML or Neurips, then senior authors will either 1) Remove themselves from some of their students’ papers (unlikely); or 2) Only submit very high quality, extremely collaborative papers which make large contributions. Point 2 will implicitly happen since senior authors are optimizing for number of papers / impact, and things will trickle down from there.
Concluding Thoughts
Although I was / am pessimistic in the short term as to how researchers will be evaluated (given that I had just finished my PhD), I have hope that the community will come together to face this challenge head on.
I will share something interesting that I have noticed after switching mainly to industry-style work over the past few months. AI has already “solved” coding in the sense that if you tell the model to do something sufficiently scoped, it will often one- or two-shot it. However, I find myself actually collaborating and discussing with my colleagues more because of this. We spend more time thinking about what we should build, and how it should be designed or work, rather than coding on our own. We will often have a lot of coding done autonomously overnight, while discussing what we should be doing during the day.
I imagine that TCS or theoretical research more broadly may become more and more like this. We may spend more time discussing results, understanding what is interesting and what we should be thinking about. The proofs may just become something of secondary import. What really matters is selecting problems well, convincing and communicating to others that a problem is important, and thinking more clearly about their impact.
Determining the difference between a regular and good paper is probably one crux of what the TCS / theory research communities will need to address in the next few years, and hence is out of the scope of this piece.
The bandits literature always gets a lot of hate. Sorry guys! Was just the first example I thought of.
If you end up needing K papers to graduate, you can certainly bang them out in a jiffy. For example, my co-authors and I all received an email from an 18 year old who claimed to have proven an open problem we posed at COLT in 2024. Upon further investigation, it appears that GPT 5.6 can one-shot the problem now. Will the student prompting GPT obtain a paper out of this? If so, this is a clear instance of the depreciation of the academic paper, and demonstrates why it surely cannot hold as the unit of currency that anyone cares about (at least for theoretical research).






