AI Agents Can Already Do Research. Now What?
What happened when we tried it ourselves, and why our institutions are not ready
Sergey Gusev and David E. Bernal Neira · Purdue University · September 2026
What the agents produced (as of 25 September 2026) · References
This summer brought two headline results in mathematics. In August, an Anthropic staff member with no mathematical training asked an unreleased Claude model to “take a real stab” at the Riemann hypothesis, one of the most famous unsolved problems in mathematics. The hypothesis says that certain zeros of a function all lie on one line. The model did not prove that, but it proved that more than 66% of them do, up from the previous best of about 42%, a share that mathematicians had raised only in small steps for decades [1]. In September, OpenAI announced that about ten thousand artificial intelligence (AI) agents, working for 88 hours, had produced a proof for a version of another famous open problem, about the Navier–Stokes equations of fluid flow; independent review is still under way [2]. By one outside estimate, that run would have cost an external customer several million dollars [3]. OpenAI has since said that the model behind that run has resolved more than 100 further open problems; it has not yet published them [4].
Both came from inside AI companies, with unreleased models and budgets that no research group has. Can a research group outside those companies, with models anyone can buy, get agents to do real research?
We tried it ourselves. It works, on a smaller scale than the AI companies’ runs and at a small fraction of their cost.
Over the past few weeks, we gave groups of AI agents a research area, anywhere from a single topic to a whole field. We gave them papers and computing tools, and told them to make real, correct, useful progress and not to stop. We gave them no additional research ideas.
Within the first day of one run, the agents had written more notes than we could read in a week. Within a few weeks, they had produced material for 45 potential scientific papers in the five areas we gave them: mixed-integer nonlinear programming, quantum interior-point methods, molecular thermodynamics, transport theory, and aggregation kinetics. In a sixth area, heterogeneous catalysis, we asked for experiments instead, and they designed laboratory experiments; nobody has run them yet, so we cannot yet say how good they are. We have not checked all of it, and we cannot yet say how much is correct or new. But in everything we have checked so far, we have found no error that would invalidate a result, and some results are proved in Lean, software in which a computer checks every step of a proof.
This post explains what we did, why we think it matters, and what we think should happen next.
The full paper has the details, the evidence, and the caveats.
All of the agents’ output is public at https://github.com/SECQUOIA/agent-swarm-research [5]; this post describes it at commit 84c6be7, before any scientific input from us.
What we did
The whole procedure fits in three steps (Figure 1).
-
Describe the problem.
Give the swarm a problem or field, enough context to start, and freedom to explore any idea.
-
Equip and start the swarm.
Provide papers and tools, with permission to get more. Ask the agents to make progress and not stop.
-
Review, write, and verify.
Have a fresh agent review every result. Write it up. Verify by proof or computation wherever possible.
We used off-the-shelf tools: the coding agents Claude Code and Codex, mostly running Claude Fable 5.1 and GPT-6 Astra, on two Claude and two ChatGPT subscriptions at $200 a month each. At pay-per-use prices, the same usage would have cost more than $15,000 in the first four weeks. We wrote no custom software beyond a few instructions for housekeeping. Anyone with access to these models can repeat this today.
Our procedure is not the best way to do research with agents: it is brute force, and we deliberately left it untuned. But if something this simple already works, better methods will do more.
We did not expect this to work. For months, we had built tools for close collaboration between people and AI, in which we directed the research and checked the work at every step. They never produced much. When we stepped aside and let the agents do everything, from choosing research directions to checking results, we had results within a day, and a lot of them. In hindsight, the reason is that in theoretical research, agents are getting better at every stage of the research process, and they are already good enough at most of them. They can direct the research and check the work themselves, and they read more than human researchers, compute more than human researchers, and explore several ideas at once. As the models improve, we can add less and less to solving a given problem, and each time we step in, we mostly slow the work down.
What came back
What came back is not a set of breakthroughs. It is incremental work, the kind of small, sound steps that make up most scientific progress. You do not have to take our word for it. Everything the agents produced is public, and anyone can check it. The tables at the end of this post list every potential paper and proposed experimental program, with what the agents claim and links to their drafts. The results from AI companies are breakthroughs on famous problems. Our runs show the everyday work of research, in several unrelated fields.
Some of this work did not need the newest models. Some of our results, in quantum interior-point methods, came from GPT-5.6 Sol, a model that had been public for almost two months before we tried it. Models could do this work before we noticed, and we suspect the same is true in many other areas of theoretical research.
The bottleneck has moved
The agents produced results faster than we could review them. At a few days for each potential paper or experimental program, reviewing everything would take one of us more than seven months of full-time work. Producing the results took weeks.
Every publication system assumes that the author has checked the work and answers for it, and that reviewers then take a second look. When one person can produce more than they can review, submitting that work directly for publication would bypass the author’s own check, and no amount of outside review replaces it. Academic hiring, promotion, and funding decisions still rely on paper counts, and a paper can now be produced for almost nothing.
If the output were noise, anyone could ignore it, and the review backlog would not matter. But the output deserves review.
For skeptics
These are some of the responses we often hear.
“The results are not impressive. A good researcher would have found them.” Maybe. If so, why were they not already in the literature? Whether agents outperform the best researcher in a field is only part of the question. How many researchers must agents be able to match before the field’s institutions need to change? Everyone except the top one percent? Everyone except one person? We think agents already match many working researchers, including us.
“It is a statistical machine. It can only recombine what it has seen.” Much of research recombines too: a known technique applied to a new problem, or ideas from two fields joined together. Whether a result is correct, new, and useful can be checked without knowing who or what produced it.
“It might all be wrong.” Some of it may be. But in what we have checked so far, we have found no scientific error that would invalidate a result. And we released all of it, so anyone can check it for themselves.
“I refuse to use these tools on principle.” That is a legitimate choice. But others in your field will use them, and you will be compared with them in hiring, funding, and promotion.
When a study shows something AI cannot do, check which model it tested and when. Such limits have repeatedly been overtaken within months, on a trend that has itself been measured [6]. If you can do better than the models today, that tells you only about today’s models, and the models keep improving.
For enthusiasts
“Research just got easier. With the newest models and enough computation, anyone can produce results.” That may well be true, but it raises problems that we cannot answer. If results depend only on the model and the budget, then the researcher is not needed. Anyone who cares about a problem can pay for the computation directly, and the company that runs the model has the newest models first and pays the least for computation. And if results follow budgets, who gets to do research, what should the human contribution be, and what is it worth?
“Then I will be the reviewer.” That is not a safe role either. We could not review fast enough. Every error we know of in our corpus was found and fixed by agents before any of us looked at the work; we cannot say whether we would have caught those errors ourselves. Agents can be run as reviewers in large numbers as easily as they are run to produce results. And nobody chooses a research career to sign off on a machine’s output. A system built on human sign-off would fill with reviews that were never really done.
So what, exactly, is the researcher’s role? What would someone pay a person for that they cannot get from the model directly? We do not have a satisfying answer.
The questions, and where we stand
We would rather take positions and be wrong than only ask questions. This is where we stand today.
-
How should we trust results when review, not production, is scarce? By measurement. Let’s put results in a form that can be checked automatically wherever possible, and build machine-checked libraries of established results in every field whose reasoning is mathematical, such as physics, chemistry, and engineering, not only in mathematics itself. Let’s have experts check a sample of what agent reviewers approve, publish how often they miss errors, and accept results from a process that meets the field’s standard.
-
Which results need a person to understand them? Let’s decide on purpose, by what is at stake and by measured error rates, and revisit the decision with each new model.
-
What counts as new when agents cannot read much of the literature, because it is behind paywalls? Let’s open the literature. Work presented as public knowledge should be readable by every agent, human or artificial, that does research.
-
What should publication become? Let’s release claims with their verification status (machine-checked, reviewed by agents, audited by sampling, or reviewed by a person) and with links to the checks.
-
What do credit and paper counts mean when the human input is a prompt? Let’s stop counting papers. Let’s judge results by whether they are correct, new, useful, and checked, and avoid punishing people for saying how they were produced. Rules against using AI cannot be enforced; they only push its use into hiding. The current system of credit is already being broken in private, and we think it is better to break it in public, where its replacement can be discussed.
-
Why would anyone pay a researcher? Less for producing results, and more for understanding them, checking them, and deciding where the work should go. There may also be more demand for researchers, not less, through an effect known as the Jevons paradox: making a resource cheaper to use can increase how much of it is used. In the 19th century, the economist William Stanley Jevons observed that when steam engines began to use coal more efficiently, total coal use rose, because cheaper power made new uses worthwhile [7]. Cheap results could likewise make far more questions worth asking and far more directions worth exploring, and people could be needed to choose among them, steer the agents, and put the results to use.
-
Does the best researcher become the one with the largest budget? Results will increasingly depend on how much computation a researcher can pay for, and we do not think that can be avoided. Let’s decide openly how to handle it, instead of drifting into it: whether institutions should provide computation as a shared resource, like libraries and laboratories, and whether assessment can separate what a person did from what their budget did.
-
How should we learn a field that machines can already work in? Let’s teach more, not less, but differently. Let’s aim training at understanding and judgment, and stop drilling students in skills that they now only need to understand. Let’s fund the training of junior researchers on purpose. On a fixed budget, one senior researcher plus computation produces more than a group of students; a field that funds research on that basis will have no senior researchers in a generation.
-
What happens when agents are run over whole fields? It will happen. Model developers, funders, and governments can afford it, and nobody can stop everyone else from doing it. We think we should do this in the open, funding checking alongside production. Running agents over a whole field also gives that field a baseline: what agents can do on their own. What people add on top of that baseline is the human contribution.
-
How do we stay in control of research we cannot keep up with? We expect agents to move far ahead of people in some fields, probably in mathematics first. Keeping up is not a realistic goal; staying in control is. That means we set the goals and can stop the runs, enough people understand each field well enough to audit it, and there are guardrails where a correct result can do harm.
What we do not know
Two of these answers are weaker than we would like.
We argue for more people, not fewer. We can say why a field needs them. We cannot yet say what they will do. Auditing, steering, and keeping up with results are the roles we can name today, and agents may take over each of them as the models improve.
And the baseline idea assumes that people can follow the baseline. If agents produce more in a month than a field can read in a year, people may not even be able to tell what is already known, let alone add to it. We do not know what the human contribution is then, or how anyone would measure it. We think it is one of the most important questions, and we raise it without an answer.
Our institutions are not ready
Peer review, paper counts, hiring, promotion, funding, and graduate training were all built for a world in which producing a result is slow and its author has read it. That is no longer the world we are in. We believe our research institutions, and the incentives they create, are not ready for what these systems can already do, let alone for what comes next. The best time to start changing them was yesterday. The second-best time is now.
Institutions change over years, and models improve over months. A curriculum or a review policy designed for today’s models will take effect after those models have been replaced, and their replacements will be replaced in turn. Such policies should be designed for capabilities that keep increasing, not for any one generation of models. Institutions should also plan ahead and decide now how they will respond when agent review is shown to be as reliable as expert review, when agents overtake people in their field, or when automated laboratories make experiments cheap to run. Some AI developers already make commitments of this kind for their own models, tying required safeguards to measured capabilities [8]. Research institutions must do the same.
Some responses have begun. More than two dozen Fields medalists, winners of the top prize in mathematics, signed a declaration warning that the push by AI companies to solve mathematical problems as benchmarks harms mathematics, and calling for the problem to be addressed urgently [9]. OpenAI has formed an independent advisory group of mathematicians to advise it on how to assess and communicate new results [4]. We welcome both as first steps, and we share the concern that results announced in a rush can get ahead of understanding. We do not think that aiming AI at famous problems is itself the harm. Anyone can now aim AI at famous problems, so such attempts will happen whether or not anyone approves; what matters is how we prepare for them. And both responses concern mathematics and the AI companies. Our runs were in optimization, thermodynamics, transport, and catalysis, with models anyone can buy. The same questions arise wherever agents can produce results faster than people can review them, and that is already true well beyond mathematics.
What to do now
-
Run the experiment in your own field. The paper includes examples of our prompts. Use the best models available when you read this. If you want to compare with us, supply no ideas of your own.
-
Report what you find: what you asked, what came back, what you checked, and what you did not.
-
Judge the results against the trend, not only against today’s models. If the agents fail, report that too, with the model and the date, and repeat the experiment as new models arrive. Capability is uneven, so the same agents may fail on one problem and do remarkable work on the next. The telling comparison is with what agents could do six months or a year ago. Then extend that trend a few years, and ask what your field should be doing now.
-
Start preparing now. Do not wait until AI clears some higher bar, such as producing results better than anyone else can. Discuss with your colleagues, students, institution, and funders what your field should change: how results are reviewed and credited, how students are trained, and what researchers are for. Prepare for capabilities that keep increasing, not only for today’s models.
The sooner many people report, the sooner this conversation rests on evidence instead of opinion. But preparation should not wait for that evidence: it should have started yesterday, and it should start now.
What the agents produced
These tables contain the same research inventory as the paper, grouped by area. Each row is a group of related results that could stand as one paper, whether or not the paper has been written. The contributions are the agents’ claims, as stated in their own write-ups and notes; we have not checked all of them. We left out results that the agents themselves marked as superseded, refuted, or withdrawn, results too small to stand as a paper, and clusters whose main results the agents’ own literature checks found to be already known. In catalysis, the agents proposed experimental programs, and none of the experiments have been run.
The links in the Write-up column go to the drafts at commit 84c6be7 (tag paper-v1), which holds the agents’ output as of 25 September 2026, before any scientific input from us; the Lean column gives the status on that date.
Later versions of the repository may add results and our own scientific input.
Rows marked “Notes only” have no draft; their links open the relevant area’s collection of agents’ notes at the same commit.
Mixed-integer nonlinear programming
| ID | Working title and claimed contribution | Write-up | Lean |
|---|---|---|---|
| M1 | Sharp gaps for positive multilinear relaxations. Disproves the Luedtke–Namazifar–Linderoth conjecture; exact hull gap; worst ratio grows as \(\ln d/\ln\ln d\) in degree and \(\ln n/\ln\ln n\) in dimension. Builds on [10]. | Full paper | Done |
| M2 | Verified bounds for positive cubic relaxation gaps. Brackets the worst cubic termwise-to-hull gap ratio: \(1610000/743033 \le R(3) \le 31/12\). | Full paper | Done |
| M3 | Checkable lower bounds for convex mixed-integer nonlinear optimization through rational outer approximations. Rational certificates give independently checkable lower bounds; 203 of 289 MINLPLib models pass separate replay; exact audits find invalid proofs accepted by an external proof checker. Builds on [11]. | Full paper | Done |
| M4 | Convex relaxation gaps and spatial certificates in nonlinear optimization. Signed bilinear gaps of the order of the square root of the edge density; exponentially many certified regions for spatial branch-and-bound under specified node oracles. Answers a question of Altschuler and Boix-Adserà negatively, assuming P\({}\neq{}\)NP; a counterexample to a published theorem on P-split formulations. Also restates the positive multilinear and cubic gap results of M1 and M2. | Full paper | Partial |
| M5 | Integer dimension in convex mixed-integer approximation of nonlinear graphs. For a fixed quadratic system on a box, the minimum number of integer or binary variables for \(\varepsilon\)-accurate convex lifts is \(\frac{1}{2}\,\mathrm{ncrank}\cdot\log_2(1/\varepsilon)+O(1)\), where ncrank is the noncommutative rank of the Hessian space; polynomial-time rational constructions. Builds on [12], [13]. | Full paper | Partial |
| M6 | Rounding switching controls under a hard switch budget: sharp minimax bounds and exact algorithms. Exact minimax rounding error for up to three switches; disproves a conjecture of Sager and Zeile and corrects a published lower bound; exact finite-grid values and algorithms. | Full paper | Partial |
| M7 | Sparse convex hulls for network flows coupled to a simplex. An exact formulation uses one extra coordinate per independent cycle of each unobserved block–label subgraph; exponential coefficient growth on series–parallel graphs. | Full paper | Partial |
| M8 | Topology, uncertainty, and precision in passive potential-flow optimization. Polynomial additive optimization of potential differences at fixed block cycle rank, for balanced nomination boxes and independent coefficient intervals; exact pressure comparison on single-source, single-sink cacti is equivalent to Square-Root Sum. | Full paper | Partial |
| M9 | The complexity of pooling: algebraic barriers and structural algorithms. The pooling threshold decision is \(\exists\mathbb{R}\)-complete; strongly NP-complete with all layer degrees two, and NP-complete with two pools and two products, answering questions of Boland et al. and Haugland; resolves a rank-one cost conjecture; matching tractability boundaries. Builds on [14], [15], [16]. | Full paper | Not applicable |
| M10 | Exact feasibility of resistive and AC power networks. Resistive power-flow feasibility is \(\exists\mathbb{R}\)-complete even on planar, degree-three, unit-conductance networks; the hardness transfers to AC networks with resistive lines. Related work: [17], [18]. | Full paper | Not applicable |
| M11 | Structured bilevel optimization with many follower variables: global responses, accuracy, and structural boundaries. A fixed-dimensional description of all global follower responses gives exact polynomial algorithms; rational leaders with error \(2^{-B}\) for strictly convex costs; hardness with dense near-identity Hessians. | Full paper | Possible |
| M12 | Globally certified measurement selection with correlated errors. Polynomial approximation sets in relative PSD order, a weighted-trace approximation scheme, and rational log-determinant certificates for correlated measurement selection. | Full paper | Partial |
| M13 | Radial and point separation for perspective outer approximation of convex generalized disjunctive programs. Matched comparison of two separation policies: radial search gives modest, implementation-specific gains; no new cut family. | Full paper | Not applicable |
| M14 | Quadratic aggregation: certificates, finite descriptions, and approximation. Resolves three Blekherman–Dey–Sun conjectures: under hidden hyperplane convexity, a hull is proper exactly when a nonconstant convex aggregation exists; a three-inequality hull needs uncountably many aggregations; four aggregations suffice for three strict quadratics with a positive definite combination, extending a Blekherman–Dunbar bound, and are sharp. Builds on [19]. | Full paper | Partial |
| M15 | Exact convex hulls for a reciprocal factor shared by many variables. Explicit hull of \((X,1/X,Y_i,XY_i)\) for many leaves, with exact rational separation and decomposition; extends to an integer factor with a binary-encoded range, with separation polynomial in the encoding length. | Notes only | Done |
| M16 | Ill-posed heat-exchanger network instances in MINLPLib.
The heatexch_gen models share an unbounded guarded log-mean-temperature term; for heatexch_gen1, a numerically feasible point improves the listed value by about 30%, and numerical estimates indicate that the infimum is not attained. |
Notes only | Not applicable |
| M17 | Convex envelopes of two-variable monomials with real exponents on a wedge. Extends the hull of a bounded monomial on a wedge from positive exponents to negative and mixed-sign ones, including ratio terms; answers Belotti’s open case of two negative exponents. Builds on [20]. | Notes only | Possible |
| M18 | Exact indicator quadratic optimization at low treewidth. NP-hard at bandwidth two with Hessians arbitrarily close to the identity; with randomly perturbed indicator penalties, exact algorithms in polynomial expected bit time at fixed treewidth. | Notes only | Not applicable |
| M19 | Convex envelopes of univariate functions of a linear form. For any lower semicontinuous \(\sigma\), the convex envelope of \(\sigma(a^\top x+b)\) over a box is a minimum over laws below a comonotone staircase law in convex order, and dually a supremum over concave minorants, each of which gives cuts valid on the whole box; extensions to products of simplices and sign-restricted order polytopes. Builds on [21]. | Notes only | Possible |
| M20 | Contraction theory of iterated optimality-based bound tightening. Near a minimizer, iterated bound tightening follows a monotone, positively homogeneous map on box shapes whose Collatz–Wielandt-type constant bounds the local linear rate; a certificate for when tightening stalls; the exact rate \((\sqrt{2a^2+4a}-a)/2\) of simultaneous rounds for \(x^2+y^2+axy\), \(0<a<2\), with McCormick relaxations; a condition under which tightening does nothing, even for strongly convex objectives. Builds on [22]. | Notes only | Possible |
| M21 | Joint relaxation of several nonlinear terms in one variable. Automatic, certified treatment in a solver: certified curvature and envelope cuts valid for every slope solve all 24 separable quartic test programs, of which native SCIP, Gurobi, and BARON solve 2, 3, and 6; a certified separator for hulls of curves \((t,f_1(t),\dots,f_k(t))\). The hull theory itself is known. Builds on [23]. | Notes only | Not applicable |
| M22 | Separable concave terms on few linear rows. A binary reformulation that allows at most \(\mathrm{rank}(A)\) variables strictly inside their concave pieces is solved in \(2n+1\) branch-and-cut nodes on a family where spatial branch-and-bound needs exponentially many (M4); the joint hull of concave terms on one row, which for equal widths is the Padberg–Van Roy–Wolsey flow-cover polyhedron in other variables, extended to inequality rows and to indicators with concave costs. Builds on [24]. | Notes only | Possible |
Quantum interior-point methods
| ID | Working title and claimed contribution | Write-up | Lean |
|---|---|---|---|
| Q1 | Objective sublevels and central-path Hessian conditioning. For every self-concordant barrier, conditioning is \(\Theta((\mathrm{diam}\,L(g)/g)^2)\); LP spectra have two scales. | Full paper | Partial |
| Q2 | The cost of following the central path. Sharp central-path movement tax \(\Gamma_r=\Theta(\sqrt{\log r})\) for objectives of rank \(r\); on an explicit sparse LP, primal–dual completion requires \(\Theta(r^{3/2})\) bounded steps. Builds on [25]. | Full paper | Partial |
| Q3 | Access models and right-hand-side mass in Newton solves. Degenerate-LP Hessians have two eigenvalue clusters that conjugate gradients handles in polylogarithmic iterations, while plain block access keeps the known \(\tilde\Theta(\kappa)\) cost; the right-hand side’s coupling to the small eigenvalues decides when filtering helps. | Summary document | Partial |
| Q4 | Winner-take-all condensation in block log-determinant SDPs. Winner mass controls conditioning; tight holonomy value queries, \(\Theta(N\sqrt{G})\) quantum versus \(\Theta(NG)\) randomized. | Summary document | Possible |
| Q5 | Accuracy curves for hidden-block state conversion. Tight accuracy-dependent query curves for quantum state conversion on proved domains; notes extend the lower bound to any inner predicate, an open problem of the summary document, and give the exact zero-query error in the range it leaves open. | Summary document | Not applicable |
| Q6 | Condition-one linear programs with hard loading and recovery. Sparse LPs with condition-one reduced Newton systems still need linearly many queries to load and recover states. | Summary document | Not applicable |
| Q7 | Limits of parity-gadget lower-bound constructions. Algebraic identities rule out natural classes of parity-gadget lower-bound constructions; notes rule out bounded local full-KKT parity amplifiers, an open problem of the summary document. | Summary document | Partial |
| Q8 | Loading and recovery hardness for semidefinite programs. Trace-normalized SDPs with condition-one Newton steps still have parity-hard values and solution states; notes give genuinely nonabelian \(\Theta(N)\) word hardness with a sparse SDP transfer, an open problem of the summary document. | Summary document | Not applicable |
| Q9 | Trade-offs between preconditioning and state interfaces. A better preconditioned condition number is paid for in normalization, recovery sensitivity, or state preparation. | Summary document | Partial |
| Q10 | Conditional speedups for sparse quantum interior-point methods. Conditional per-step gains from certified refresh, face repair, compressed dual output, and block-angular borders. | Summary document | Partial |
| Q11 | A correction to the complexity analysis of the quantum central path method. Both the simulator-norm bound and the clock analysis of arXiv:2311.03977v2 fail as claimed; corrected norm and speed trade-off. The authors informed us that they already knew of at least one error and are preparing a revision; no correction was public when the agents ran. | Summary document | Done |
| Q12 | Curvature, support certificates, and barrier complexity of conic lifts. An exact lift by definable cones of a body with a strictly curved boundary patch needs \(\sum_i\max(\dim K_i-2,0)\ge s-1\); the least support-certificate rank of an \(s\)-dimensional ball is exactly \(\lceil (s-1)/B\rceil\); exact barrier parameters for symmetric-cone balance slices. Also: disk-product instances whose central states need \(O(1)\) queries but whose scalar readout needs \(\Theta(N)\); with matched oracles at the fiber center, PSD packing gives no query advantage for reduced Newton solves; compilation of sparse second-order cone constraints to barrier parameter at most \(k+1\) with \(\Theta(\sqrt{Nk})\) quantum versus \(\Theta(N)\) randomized queries. | Full paper | Possible |
| Q13 | Classical and quantum query complexity of scalar Newton quantities. Relative estimation of \(b^*H^{-1}b\) with \(\kappa\varepsilon^{-2}(d+1)^{O(\sqrt\kappa\log(2/\varepsilon))}\) classical queries and a lower bound matching the known \(\tilde O(\alpha\kappa/\varepsilon)\) block-access algorithm up to logarithmic factors on \(3\times3\) matrices; statistical-query sampling of Newton steps; one-cone sparse SOCP: \(O(1)\) quantum queries versus \(\tilde\Omega(N^{1-1/k})\) classical statistical queries. | Full paper | Not applicable |
| Q14 | The quantum cost of unit-normalized spectral shifting. Answers an open problem of the summary document: a sharp query staircase \(\Theta(\delta^{-1+1/(2\ell)})\) for block encodings of \(I-H\) with normalization one, and \(\Theta(\delta^{-1}\log(1/K))\) at high accuracy; a sparse LP separates normal-matrix from factor access. | Full paper | Not applicable |
| Q15 | Exponential-cone scenario compression for entropic risk. Compresses \(N\) scenarios into \(O(1+K+\log)\) exponential cones, independent of \(N\), with certified value bounds; matched \(\Theta(e^K/\varepsilon)\) quantum versus \(\Theta(e^{2K}/\varepsilon^2)\) classical source queries under the stated access model. | Notes only | Not applicable |
Molecular thermodynamics
| ID | Working title and claimed contribution | Write-up | Lean |
|---|---|---|---|
| TD1 | Finite reservoirs at phase coexistence: full-state accuracy and phase correlations. When finite baths reproduce canonical laws; an \(N^{3/2}\) threshold for the two-phase systems studied under stated phase-tail conditions; a shared bath of intermediate size drives two copies into opposite phases while each copy alone stays canonical. Builds on [26]. | Full paper | Not applicable |
| TD2 | Survival-conditioned thermodynamic integration. Endpoint-survivor force integration is path-dependent, with error of second order in weak killing; interior survivors admit an exact potential. | Notes only | Possible |
| TD3 | Interfacial tension from bulk response in nonlocal double-parabola models. Exact tension certificates from bulk response; fourth-order moment matching leaves the tension undetermined. | Notes only | Not applicable |
| TD4 | Capacity certificates for reversible nucleation kinetics. Capacity lower bounds from conditional transport; paired finite-field bounds on nucleation-rate response. | Notes only | Not applicable |
Transport theory
| ID | Working title and claimed contribution | Write-up | Lean |
|---|---|---|---|
| TP1 | Designing surface transport under uncertain kinetics: moment thresholds and measurement precision. For a small surface-mobility budget \(M\), the best fixed placement raises the disorder-moment order at which rare merging defects dominate from \(4/3\) to \(8/5\); observing the defects improves the mean from \(M^{-1/4}\) to \(M^{-1/5}\), and resolution of order \(M^{1/5}\) suffices. Builds on [27]. | Full paper | Not applicable |
| TP2 | Kinetic defects in adsorbing channels. Weak surface diffusion \(D_s\) at a quadratic kinetic minimum gives a \(D_s^{-1/4}\) dispersion divergence with an explicit crossover to a rate floor; when the minimum’s location is known, optimal placement of a mobility budget \(M\) improves the divergence from \(M^{-1/4}\) to \(M^{-1/5}\), the known-defect case behind TP1; an exactly solvable placement transition. | Notes only | Not applicable |
Aggregation kinetics
| ID | Working title and claimed contribution | Write-up | Lean |
|---|---|---|---|
| AK1 | Sampling-law separation and finite nonlinear corrections in additive coagulation–fragmentation. Sharp number–mass separation bounds, a finite log-size correction, finite-population breakdown, Fourier identification. Builds on [28]. | Full paper | Not applicable |
| AK2 | Survival under unobserved sister-type dependence in multitype branching. Uniform near-critical error bounds for extinction over unobserved daughter couplings, with an exact optimal coupling and a rule for tied reproductive values; the basic sharp envelopes follow from known branching and rearrangement results. | Notes only | Possible |
Heterogeneous catalysis: proposed experimental programs
| ID | Working title and claimed contribution | Write-up | Lean |
|---|---|---|---|
| CA1 | Physical water management in Fischer–Tropsch synthesis. Tests whether late hydrophobic-polymer addition protects conditioned cobalt; re-analysis of published data finds a product output about 1.9 times that of the reference with the polymer. | Program document | Not applicable |
| CA2 | Steam compatibility of cyclic oxides in chemical looping. Tests whether steam purges, assumed by a published process model but not tested, change ethylene output from coated LSF; compares CO\(_2\) supplied with the steam (protection) and afterwards (recovery). | Program document | Not applicable |
| CA3 | Catalyst demand in polymer ethenolysis. Tests whether lower ethylene pressure can replace part of the fresh Na/alumina catalyst when reused catalyst converts polyethylene to propylene. | Program document | Not applicable |
| CA4 | Nickel and the useful life of promoted silver epoxidation catalysts. Tests whether nickel adds ethylene oxide output beyond chloride policies; a 1995 patent already reports a retention benefit. | Program document | Not applicable |
| CA5 | Tungsten coordination and retention in sugar conversion. Tests whether a controllable anchoring or feed variable links productive sugar coordination, tungsten loss, and glycol output. | Notes only | Not applicable |
| CA6 | Product-rich liquid Ti-zeolite epoxidation. Tests whether a reaction network calibrated on dilute kinetics predicts epoxide output and peroxide loss once products accumulate. | Notes only | Not applicable |
| CA7 | Oxygen fate and self-cleaning in zirconia-catalyzed styrene production. Tests whether an oxygen-removal pathway predicts sustained styrene output and a feed policy. | Notes only | Not applicable |
| CA8 | Acid-site assays and zeolite aging. Tests whether repeated NH\(_3\) and water assays change the hydrothermal aging of H-CHA. | Notes only | Not applicable |
Status values:
-
Write-up. Full paper: a complete, compiled manuscript. Summary document: part of the long document that collects the quantum interior-point results, which we asked for instead of separate papers. Program document: a chapter of the document that ranks the proposed catalysis programs. Notes only: the results exist only as the agents’ notes and checks.
-
Lean. Done: the main mathematical results are proved in Lean, with no unproved steps; software and experiments are not covered. Partial: some of the results are proved in Lean, but not all. Possible: not done, but the main claims are mathematical statements that could be formalized with current libraries; this is a judgment, not a check. Not applicable: the main claims rest on numerical evidence, experiments, or modeling assumptions, or require a framework that current formal libraries do not provide, such as complexity classes, quantum query models, or limit theorems for stochastic processes.