Memory-Efficient Optimizer States for Multi-Agent Financial LLM Systems
Source: AegisMind Research
Read full discoveryAutonomous Scientific Discovery
solver.press is the public window into AegisMind's discovery engine. Hypotheses are generated, stress-tested through five-model adversarial debate, checked for logical consistency, and turned into experimental packages designed to be cheap to kill.
Everything here is a research lead. No hypothesis on this site has been validated in a wet lab. Where computation is itself the experiment — ML theory, quantum systems, mathematics — the numerical results stand on their own. In biology they do not, and we label them accordingly.
Discoveries that have had a computation run against them, most recent first. Where that computation is itself the experiment it is marked verified; where it is a structure-prediction score against a protein target it is marked as docking only, which is a screening result and not evidence of binding.
Source: AegisMind Research
Read full discoverySource: AegisMind Research
Read full discoverySource: AegisMind Research
Read full discoveryHighest-confidence discovery
Medicine · Efflux pump inhibitors targeting MexAB-OprM restore carbapenem susceptibility in MDR Pseudomonas aeruginosa by reducing …
Read full discovery →Running since March 2026. These are research leads that survived adversarial debate — not findings, and not confirmed.
Source: AegisMind Research
Read full discoverySource: AegisMind Research
Read full discoverySource: AegisMind Research
Read full discoverySource: AegisMind Research
Read full discoverySource: AegisMind Research
Read full discoverySource: AegisMind Research
Read full discoveryThe funnel, honestly
2,460
hypotheses generated
Since March 2026. Read live from the discovery database on each page build, not a snapshot.
142
visible on this site
Survived debate scoring. Visible means published here — not verified. Withdrawals are counted, not hidden: 26 on 7 August 2026 when a parsing bug turned fragments of a model's reply into hypotheses; 8 on 24–25 August, seven built on multiple-sclerosis targets that did not replicate and one on a scaling law measured on a model that never trained; 8 on 30 August whose adversarial critic had found a category error; and 10 more the same day — three that were not hypotheses at all but internal assessments published by mistake, three resting on a scaling law our own paper has since retracted, one proposing to reduce by 30% a barrier we now measure at zero, one applying a neural-network quantity to electrolyser hardware where it has no definition, and two built on a resistance-gene network that is synthetic — in which the gene named as the hub turns out not to be a bridging gene at all. An eleventh was withdrawn and then restored the same day: the reason recorded against it did not hold. A further 8 published records are NOT listed here, for lack of an experimental-validation package — reachable by direct link only. The gap is counted rather than papered over.
6
where computation settles it
A negative counts. Computation settling a question against us is computation settling it — 3 refuted, 3 supported, across two unrelated fields. Refuted: 72 compounds scored against a bacterial efflux pump and an unrelated human protease correlate at r = +0.932, so a docking-score gap between two proteins is not selectivity; and an entropic-transport hypothesis that proved not to be independent of the one beside it — its only distinguishing prediction, a slower rate for non-smooth problems, measures 1.0004 against a smooth control at 1.0018. Supported: an O(ε) convergence rate that holds unchanged to 1,000 dimensions, and — after correcting two of our own claims — that what is computed from docked poses moves the answer while how hard they are searched does not. Of the rest, 1 were inconclusive, 76 are literature-tier — a model reading papers, which settles nothing — and 67 have no computational field at all. Benchmark runs where computation IS the experiment are not discovery records and are not counted here.
We show this because the ratio is the most honest thing we can tell you about the system. Generating plausible, well-argued, internally consistent hypotheses at volume is demonstrably easy. Establishing that any one of them is true is not, and we have not yet done it for a single biological claim.
Wet-lab confirmation is deliberately not a column here. This funnel counts stages our own process runs, and that is not one of them — an assay is somebody else’s laboratory deciding to spend its time on us. Until 11 September 2026 a fourth card carried a hard-coded zero for it, which reported a third party’s calendar as though it were our result. The fact it stood for has not changed and is not hidden: no hypothesis on this site has independent experimental confirmation. When a verdict exists, it will be shown as a verdict — either way — rather than as a zero.
Track record
0.4530
AutoDock Vina on SARS-CoV-2 Mpro, 751 measured compounds, repaired receptor — below chance. Factor Xa 0.6775, PD-L1 0.5356.
0.7654
Seven descriptors — weight, clogP, donors, acceptors, rotatable bonds, TPSA, charge. Microseconds of compute, and docking adds nothing resolvable on top.
0.8588
Descriptors + Boltz-2 on the same compounds: +0.093 [+0.062, +0.125] over descriptors alone, replicated on Factor Xa at +0.075. The first thing we tested that cleared a bar fixed in advance.
The deployable number is the smaller one. That 0.8588 is fitted on the target’s own actives, so it needs answers you would not have on a new target. Boltz-2 alone — sequence and SMILES, no labels, no receptor preparation, no docking box — reaches 0.7913 on Mpro and 0.7227 on Factor Xa. Fit the descriptor model on one target and apply it cold to the other and it lands below chance, so the descriptor half is target-specific bias, not chemistry that travels.
Six interventions on Mpro, each rule fixed before its data existed
| Adding Boltz-2 to the descriptors | +0.093 | clears |
| Changing the scoring function, identical poses | +0.140 | clears — but see below |
| Scoring the pose ensemble, not the top pose | +0.055 | clears, all 40 fold seeds |
| Repairing the receptor | +0.045 | marginal against the 0.039 floor |
| Supplying the binding site to Boltz-2 | −0.013 | blind scored higher |
| Eight times the search effort | −0.019 | top ten unchanged |
| Swapping Vina's poses for DiffDock's | — | not evaluable — withdrawn |
There are two floors, not one, and we had them conflated until September. Scoring the same poses twice moves AUROC 0.020; re-running the whole protocol moves it 0.039. Against the right floor for each row, two clear, one is marginal, two fall inside, and one cannot be measured on this panel at all. The 8× search row is the sharpest: its effect is smaller than the reproducibility of the protocol it was measured in.
We tried to generalise it, and failed. One method beating six docking scores is one point, not a pattern, so we pre-registered a third scoring approach to test whether the ceiling belongs to docking or to these panels. FlashBind residualises to 0.538— inside the docking band of 0.494–0.573, not above it — with all three controls passing and a pocket check ruling out the obvious confound. Boltz-2 is still the only point outside. The claim we published as open stays open.
The second target replicates the direction, not the size. On Factor Xa the Boltz-2 residual is 0.588against that target’s own Vina arm at 0.573 — above it, but by +0.015where Mpro’s margin is +0.083, and the twelve-seed ranges touch. Our abstract used to say the two do not overlap; it now says that of Mpro only.
Tested on someone else’s benchmark.On LIT-PCBA’s 15 targets the same seven descriptors reach a median 0.726 against published AutoDock Vina at 0.581 — docking numbers we did not produce. On its AVE-debiased variant they fall to 0.651, so part of that margin is property bias the debiasing removes, and part is not.
The same method, a field away: performative optimization
Not a drug-discovery result. A convergence claim in stochastic programming, tested the same way — criteria fixed first, then look. The O(ε) rate holds, unchanged from 20 to 1,000 dimensions (α = 0.9997), which had never been shown. The bound on its constant does not: it runs from 1.05 to 2,209times the Lipschitz estimate as the problem stiffens, so a criterion published without a condition number was measuring well-conditioned problems rather than a method. The paper’s second hypothesis proved not to be independent of its first, and its one distinguishing prediction is refuted. The compute budgeted for the 1,000-dimensional work was 12 GPU hours; it took 32 seconds of CPU.
What this cost us to believe. Every receptor we had ever docked against carried zero hydrogen-bond donors — the converter typed all 754 Mpro nitrogens as acceptors, so the scorer could not form a protein-donor hydrogen bond at all. Published numbers of ours that we have since corrected: a pose-ensemble refutation that was a 2nd-percentile fold draw, a pose-generator result withdrawn as not evaluable, a Vina baseline we quoted at 0.61 that is 0.581 at its primary source, a residual band computed two different ways, and a convergence paper that reported two hypotheses confirmed when its own results file recorded otherwise. PD-L1 is closed: seven descriptors reach 0.9145 on it, so nothing docked against it bears on binding. How we decide what counts →
Ingest
arXiv and Semantic Scholar papers across 15+ scientific domains are continuously embedded into a vector store and searched for connections that span fields. In practice the engine usually lands inside established, well-populated literatures rather than on untouched ground — efflux-pump inhibition, β-lactamase inhibitor analogues and cathepsin inhibitors are all crowded fields with decades of work behind them. The useful output is a specific, narrow angle within a known field, not a bridge nobody has crossed.
Generate & Debate
Five frontier AI models (Claude, GPT, Gemini, Grok, Mistral) independently generate hypotheses then critique each other in adversarial debate. A debate score reflects survival under critique. This measures how well a claim withstands argument — which is not the same as being true, and several hypotheses that scored well here were later refuted.
Logical consistency check
Z3 checks whether a hypothesis contradicts itself — not empirical truth, but internal coherence. About 62% of hypotheses pass. A pass means the claim is not self-contradictory; it says nothing about whether the claim is true, and we do not count it as evidence.
Experimental Validation Package
Surviving discoveries receive a full EVP built around the cheapest experiment that could kill the hypothesis: precise quantitative claim, explicit disproof criteria, protocol, abort checkpoints, implementation code, compute and cost estimates, and a dependency map.
Where the loop actually closes
For ML, quantum, and mathematical claims, we run the computation on Google TPU Research Cloud and the result settles the question — the loop closes. For biological claims it does not: computation produces a prediction, and only a wet lab can close it. Those EVPs are handed off, not resolved here.
Hypothesis Aggregation Papers
View all →Formal papers synthesising solver.press discoveries into testable hypothesis clusters with complete experimental validation packages.
CLOSED 10 September 2026, and this card overstated it until then. H₁ as written is a conjunction and its second half fails: the Flory-Huggins C* ≈ 3.5 µM for Q46 is a two-point calibration fitted to Peskett et al. (2018), not a prediction from sequence, and the quantitative link to the −44.7% target gene deficit does not hold — a parameter sweep shows the predicted depletion is set by the mHTT nuclear volume fraction rather than the partition coefficients, and reaching −44.7% would need that phase to occupy 16–31% of nuclear volume against 0.4–3.8% at plausible values. What closed the line is the tissue itself. In GSE3790 HD caudate (38 HD, 32 control) the two genes that would distinguish sequestration from ordinary cell loss, BDNF and PGC-1α, do not move at any Vonsattel grade including grade 0, while the medium-spiny-neuron markers PPP1R1B and DRD2 fall exactly as neuron loss alone predicts. Those two genes were the Phase 3 readout of the EVP below, so the experiment we had costed is testing a signature that is already absent in human tissue. H₂ (BET bromodomain inhibitors dissolve mHTT condensates, restoring TF availability) was never tested and is not refuted — it is unsupported, because the mechanism it was meant to rescue has no downstream footprint to rescue. Patent AU2026905785 is retained on commercial grounds, not evidential ones.
REFUTED 20 August 2026 — the hypothesis was that strongly connected components in the collateral sensitivity graph define closed evolutionary traps. Tested against 104,337 susceptibility records from BV-BRC, a 3-node component in K. pneumoniae {imipenem, meropenem, tetracycline} satisfied every check applied: FDR correction, permutation p = 0.001, drug-class specificity against tigecycline, stability across year bands. It does not survive stratification by clonal lineage. Within MLST sequence types the association is absent (pooled OR 0.93 and 0.95; CMH p = 0.73 and 0.80) and the two dominant lineages pull in opposite directions (ST307, 1.26; ST258, 0.67); size-matched random partitions leave the odds ratio at 1.76, so the collapse is lineage-specific and not an artefact of stratifying. The tigecycline contrast, previously offered here as evidence against a clonal explanation, weakens under the same stratification. The manuscript was withdrawn from journal review and the isogenic follow-up is withdrawn with it — an isogenic experiment is a within-lineage comparison by construction, and that comparison has already been made. The E. coli edge is not pursued: 87 co-tested isolates, and it fails if a single isolate is reassigned. What survives is methodological and is reported separately: the lineage-adjustment control itself. Pooling the same contingency tables by sequence type, then permuting the sequence-type labels while preserving stratum sizes, separates genuine confounding from an artefact of stratifying at negligible cost — and the two larger published clinical collateral-sensitivity analyses were not adjusted this way. That paper is prepared and not yet submitted. Data and code for every figure here are deposited at Zenodo (10.5281/zenodo.22036517), and the preprint record retains its earlier versions at its own DOI, so the claim and its withdrawal both stay readable; a further correction is in screening.
The hypothesis: combined QS-inhibitor plus QS-dependent antibiotic therapy creates a doubly unfavourable evolutionary landscape for resistant mutants, with selection coefficient s ≤ −0.05 per passage under combined therapy and s > 0 under monotherapy. THE MODEL DOES NOT PRODUCE THIS (25 August 2026). This card previously said the Lotka-Volterra public-goods model 'predicts' s ≤ −0.05. It does not, and the paper never claimed it did — §3.1 states the model 'does not yield a precise numerical prediction because b and c_G are uncharacterised'. The −0.05 is borrowed from measured cheater costs in Sandoz et al. 2007 and chosen to clear a drift threshold. Sweeping 400,000 parameter sets, the two halves of the hypothesis are met together in 0.0% of the space — and the reason is structural rather than statistical. Combined therapy is what drives public-goods availability toward zero, s decreases in that availability, so s under monotherapy is at most s under combined therapy for every possible parameter draw; the hypothesis needs the reverse. No values of b, c_G or the antibiotic terms can satisfy both arms. In the stated fitnesses the cheater's advantage is the avoided production cost, which does not depend on public-goods availability at all, so removing the goods leaves that advantage untouched. The biology may still hold — measured cheater costs are real — but the derivation offered for it does not, and the study as designed cannot come out the way it is written to.
One hypothesis of two, and this card claimed both until 11 September 2026. H₁ holds and holds well: exact O(ε) convergence of performatively stable solutions to classical SP optima across all 5 problem families, α = 1.000–1.028 with R² ≥ 0.9995, unchanged at d = 1000 (α = 0.9997) with L̂ constant across three orders of magnitude — which had never been shown. Two things go the other way. The proportionality constant is NOT characterised as this card used to say: the criterion C ≤ 1.5·L̂ was met by 2 of 5 families, and C/L̂ runs from 1.05 at μ = 1 to 2,209 at μ = 0.05 while the problem is still convex, so a criterion published without a condition number was measuring well-conditioned problems rather than a method. And H₂ (the entropic-optimal-transport analogy) is refuted as an independent hypothesis: §2.4 relates it to H₁ by ε ↔ 1−ε, and its only distinguishing content — α < 1 for non-smooth displacement — measures 1.0004 on a discontinuous step map against 1.0018 on a smooth control. The compute budgeted for the 1,000-dimensional work was 12 GPU hours; it took 32 seconds of CPU.
SHELVED 4 August 2026 — the central claim does not survive re-testing. The reported 84.9% ergotropy improvement at large cavity detuning was measured with the qubit initialised already excited, so no energy had to be transferred: it is a retention result, not a charging protocol. Re-run with an explicit charging phase (cavity charger → qubit battery), the optimum moves to exact resonance and detuning is strictly harmful — at the claimed optimum of −10g the battery ends with zero ergotropy. The original parameters could not charge at any detuning in any case, since g/κ = 1.0 puts photon transfer (π/2g = 15.7) slower than cavity lifetime (1/κ = 10). H₂ falls with it, established 21 August: the interpolation simulation uses the identical initial state, so its 54.7% is also a retention result, and its optimum g* = 0.01 gives g/κ = 0.1 — where charging cannot occur at all. Both optima sat at the edge of their sampled grids. The simulation and statistics were sound; the initial state was not.
Preprints
3 preprints generated by the AegisMind discovery loop. These are self-deposited to Zenodo and Research Square — they have DOIs and are citable, but they have not been peer reviewedand no journal has accepted them. “Published” here means publicly deposited, nothing more.
Two earlier deposits are no longer listed here, and are named rather than dropped. The MSH3 ATPase inhibitor paper (Zenodo 10.5281/zenodo.20586369) is retracted — its target analysis used the wrong chain of PDB 3THW, which is MSH2, the same error recorded against the Precision Tetrahedron above. The physical time-lock encryption paper (Zenodo 10.5281/zenodo.20627819) belongs to a programme closed in August 2026 with no claim refuted: free incumbents already do what it proposed. Both records stay live at their DOIs so the claim and its withdrawal remain readable.
SEPTEMBER 2026 — TWO CORRECTIONS. The C* figure is not a prediction from sequence: both the Flory-Huggins χ parameter and the concentration scale are fitted to Peskett et al. (2018), and the polyQ-length dependence is imposed by the χ(Q) parameterisation rather than demonstrated by it. Separately, the mechanism's downstream signature is absent in the tissue it is claimed to occur in: in GSE3790 HD caudate (38 HD, 32 control) the positive control PPP1R1B falls 1.07 log2 (p = 5×10⁻¹¹) and DRD2 falls on all four probes, but both are medium spiny neuron markers and HD caudate loses those neurons — their decline is what cell loss alone predicts. The genes that would distinguish sequestration from cell loss, BDNF and PGC-1α, do not move (13.5th and 56.9th percentile of 22,283 probes; 46% of probes are down in HD caudate, so direction alone is uninformative). Nothing here tests partitioning. Flory-Huggins phase diagram predicts C* ≈ 3.5 µM for Q46 — physiologically accessible in HD striatal neurons. A sensitivity analysis shows the predicted TF depletion is governed by the mHTT nuclear volume fraction, which current data do not constrain, and at plausible values falls one to two orders of magnitude short of the −44.7% striatal expression deficit — so co-partitioning is presented as a mechanism to test, not as an account of that deficit. BET bromodomain inhibitors (JQ1, OTX015) hypothesised to restore TF availability by disrupting mHTT condensates. Patent AU2026905785.
Loss landscape topology across number formats and multi-target drug discovery. TWO SEPARATE RETRACTIONS. v2 (July 2026): the reported EPTIFIBATIDE dual-target KPC-3/MSH3 convergence does not hold — the MSH3 oracle used the wrong chain of PDB 3THW, which is MSH2, making the three analyses pseudo-replicates of one another. August 2026: the headline scaling law (FP32↔BF16 barrier ∝ params⁻⁰·⁴⁷, R²=0.99) is withdrawn. Its largest of three data points was measured on a model that never learned — 6.78 bits/char against a 6.02 uniform-random floor — and we reproduced that failure independently on different hardware. A barrier measured on an untrained model is not a measurement: two runs of the same failed 38M model give 0.084 and 0.409 nats, a five-fold spread. Excluding it leaves two points, which is a line, not a law. The paper's own body called this a 'directional observation, not an independently-powered scaling law'; the abstract oversold it. The KPC-3 result is also not what it appeared: the compounds lack the anionic N6-sulfooxy warhead that defines DBO β-lactamase inhibition, and the screen docked KPC-2 (PDB 3RXX), not KPC-3. What stands is the format geometry — that exponent range, not mantissa width, sets the inter-precision barrier.
A positive-control / methods-validation study. A four-phase pipeline over a 32,239-cell scVI atlas re-recovered Cathepsin S, BET/BRD3 and DNMT1 — targets already established in the literature — from public data in CA-RIM lesions. The point is that the pipeline recovers known biology unprompted, not that these targets are new. An earlier version of this work framed them as novel and used a 36,966-cell atlas that had wrongly pooled in 25 unrelated COVID-19 patients; both are corrected in the current version. A further result belongs here even though it does not touch the methods claim: when we tested these targets in an independent cohort (GSE138614), they did not hold. Cathepsin S was non-significant, and ZNF740, DNMT1 and the top discovery gene all reversed sign; of 21 genes tested only two approached significance, neither of them focal, which at 21 tests is chance. The associated patent was lapsed on that basis, and eight discoveries built on these targets were withdrawn from this site in August 2026. The pipeline did recover targets the literature already knew about — that is what the study claims — but our own replication does not support them.
solver.press publishes a curated selection of AegisMind's discoveries. Research teams, pharma BD teams, and technology organisations interested in domain-specific discovery runs can get in touch via aegismind.app.
Each EVP includes: precise quantitative hypothesis · disproof criteria · full experimental protocol · abort checkpoints · implementation code · GPU hour and cost estimates · ROI projection · prerequisite dependency map · downstream discovery unlocks.
The engine is running. The loop is closed where computation is the experiment, and open everywhere else — which is where we would want to work with you.
Visit aegismind.app →