Singapore’s AutoTrust Posts Top Marks on Public Tests of Recursive Self-Improving AI With a Sliver of Rivals’ Funding
![]() |
Its ScienceGuru system led two public benchmarks and posted the fastest reported times on two others in September
SINGAPORE, Sept. 28, 2026 /PRNewswire/ — A startup with about 10 researchers spent September setting the pace on public tests of AI that improves AI, a goal rivals have raised hundreds of millions of dollars to pursue.
AutoTrust AI said Monday that ScienceGuru, its agentic research system, posted results on four public benchmarks used to test self-improving AI: first place on Autoresearch@Home, the top validated entry on MedARC’s NanoPath when posted, and the fastest times reported to date on two GPT-2 training speedruns, both self-reported. Running on AutoTrust’s own Guru Turbo models, ScienceGuru proposed, implemented and tested each recipe; the company’s researchers set the objectives and supplied the computing power.
Recursive self-improvement, or RSI—AI that improves how AI itself is built—is now an explicit goal at the industry’s frontier. OpenAI said this month that, by its own measurements, it had reached a goal set publicly last October: an automated AI research intern. The four benchmarks ScienceGuru tackled serve as public testbeds for RSI.
“Recursive self-improvement is the race that will decide who builds the next generation of AI, and it can now be measured in public,” said Daniel Tang, AutoTrust’s co-founder and chief executive. “In September, a small team in Singapore took first place on two of those tests and posted the fastest reported times on the other two, with a fraction of the capital of the labs we are measured against.”
Four tests, one month
The four results, in the order posted:
- Autoresearch@Home (Sept. 1): First place on the official leaderboard. Built on autoresearch, a project from OpenAI founding member Andrej Karpathy, the community benchmark gives an AI agent one GPU and five minutes of training to improve a language model. ScienceGuru’s entry scored 0.8895 validation bits per byte, where lower is better, and 20 reruns of the same code with a fixed seed averaged 0.8897.
- MedARC NanoPath v2 (Sept. 9): Top validated entry when posted. The benchmark’s maintainer independently retrained ScienceGuru’s pathology-AI recipe with three fresh seeds and validated it at 0.6597, where higher is better. When AutoTrust posted the result, the official leaderboard listed it as the validated leader (snapshot at 2:17 a.m. UTC on Sept. 9). It ranked second among validated entries as of Sept. 26.
- NanoGPT Speedrun (Sept. 25): 24.90 seconds, 2.7 times as fast as the official record. ScienceGuru trained GPT-2 Small to the 3.28 validation-loss target of Keller Jordan’s community speedrun on AutoTrust’s own node of eight NVIDIA H100 GPUs, averaging 24.90 seconds over five fixed seeds with every run reported. That compares with the official record of 67.56 seconds. Self-reported; not yet reviewed by the benchmark’s maintainers.
- Time-to-GPT-2 (Sept. 25): 72.24 minutes, 27% less time than the official record. ScienceGuru trained a GPT-2-grade language model on a single 8×H100 node in 72.24 minutes, against an official record of about 99 minutes on Mr. Karpathy’s nanochat leaderboard. That beats the best community recipe’s six-run average by 9.6 minutes and the fastest community experiment’s three-run average by 1.7 minutes. By AutoTrust’s review of public submissions, it is the fastest time reported to date; it comes from a single, self-reported run.
How ScienceGuru did it
ScienceGuru ran the same research loop on every benchmark: survey the published work and open submissions, combine the strongest ideas, implement and test them on the benchmark’s standard hardware, verify results with pinned source code and complete evaluations, then publish code, verification materials and credit. All of it is at github.com/AutoTrustAI.
On the NanoGPT Speedrun, ScienceGuru fused two pending community submissions—ANVIL2 by Deven Pietrzak and Exact-match by Herman Brunborg—cut the training schedule from 1,194 steps to 652 and re-engineered how the server’s processors feed its GPUs. On Time-to-GPT-2, it started from nanochat and a community recipe by Giovanni Zinzi, narrowed the model’s feed-forward layers to about 14% below nanochat’s default width and set a 9,841-step training horizon. On Autoresearch@Home, it proposed and implemented architecture, optimizer, kernel and memory-layout changes within the fixed five-minute budget.
A fraction of the capital
AutoTrust produced the results as a seed-stage lab that has raised less than $10 million to date. Other startups pursuing automated research have raised far more: besides Recursive’s $650 million, Periodic Labs launched in 2025 with a $300 million seed round. Of those labs, Recursive is the only one AutoTrust found with published results on these benchmarks, and ScienceGuru’s figures are better on both tasks the two share: a self-reported 24.90 seconds on the NanoGPT Speedrun against the 75.36-second June record credited to Recursive’s system, and a leaderboard score of 0.8895 on the five-minute autoresearch task against the 0.9109 average Recursive published from its own runs.
From benchmarks to business
The benchmarks are a proving ground; the product is the system behind them. ScienceGuru is the top layer of what AutoTrust calls an agentic operating system for research and model development.
The stack integrates four capabilities:
- Proprietary models: the Guru family of foundation models. Guru Turbo 1.0 and 1.2 powered all four September results.
- Agentic harness: ScienceGuru, which reads the literature and prior work, forms hypotheses, writes and runs code, audits results and writes them up across long research sessions.
- Model harness: tooling that post-trains, evaluates, compresses and deploys models, including Blocks of Experts, LoRA, reinforcement learning and mixture-of-experts rewiring, plus Neural Architecture Searching, i.e pruning and quantization tuned to the target hardware.
- RSI capabilities: ScienceGuru supervises the training of Guru models alongside AutoTrust researchers, and verified research trajectories become training signal for the next generation of models.
That last layer is what makes the system recursive: ScienceGuru runs on Guru models and helps train the next ones.
“What’s unique is the combination: our own models, an agentic harness that does real research, and model and memory harnesses that turn what it learns into better models,” Mr. Tang said.
ScienceGuru is available now for macOS and Windows at scienceguru.ai, and AutoTrust uses the same stack to build customized, sovereign models that enterprises train on their own data.
“Every recipe behind these results was proposed, implemented and tested by ScienceGuru, and the code is public so anyone can check it,” said Josh Liu, AutoTrust’s co-founder and chairman. “Customers get the same loop today: ScienceGuru for their research, and the same harness to train and deploy models on their own data.”
About these results
Standings are as of the dates shown and can change as new entries arrive. The NanoGPT Speedrun and Time-to-GPT-2 results are self-reported by AutoTrust and have not yet been reviewed by the benchmarks’ maintainers. Comparisons use figures published by each source, checked Sept. 24–26, 2026. The benchmarks are independent of AutoTrust.
About AutoTrust AI
AutoTrust AI Pte. Ltd. is an applied AI research lab headquartered in Singapore, with an office in Silicon Valley. It builds the Guru family of foundation models and ScienceGuru, an agentic platform for scientific research and model development, and trains customized sovereign models for enterprises. Learn more at autotrust.ai and scienceguru.ai.





