Notes

Evidence on AI R&D Progress from NanoGPT

21 April 2026

Classifying human and agent contributions to the NanoGPT speedrun, and what publicly tracked challenges can tell us about AI R&D acceleration.

Fine-tuning experiments on CoT controllability

1 April 2026

We find that a small amount of fine-tuning on instruction following in the CoT generalizes to meaningful increases in CoT controllability on an out-of-distribution set of tasks. We fine-tune four reasoning models on small datasets of instruction-following reasoning data and OOD controllability rises from an average of 2.9% to 8.8% across four models.

Impact of modelling assumptions on time horizon results

20 March 2026

Alexander Barry examines how different modelling choices affect METR's time horizon estimates.

We spent 2 hours working in the future

19 March 2026

Thomas Kwa describes a tabletop exercise where METR researchers simulated having access to ~200-hour time horizon AIs.

Many SWE-bench-Passing PRs Would Not Be Merged into Main

10 March 2026

We find that roughly half of test-passing SWE-bench Verified PRs written by recent AI agents would not be merged into main by repo maintainers. A naive interpretation of benchmark scores may lead one to overestimate how useful agents are without more elicitation or human feedback.

Observations from two CLI game reimplementation runs with Opus 4.6

3 March 2026

Nikola Jurkovic describes observations from tasking Opus 4.6 with reimplementing Slay the Spire and Balatro in the CLI.

Five lessons from having helped run an AI-Biology RCT

19 February 2026

Luca Righetti shares takeaways on the role of randomized controlled trials in AI safety testing.

Analyzing coding agent transcripts to upper bound productivity gains from AI agents

17 February 2026

Amy Deng investigates whether coding agent transcripts could serve as an alternative for estimating AI productivity uplift, using 5305 Claude Code transcripts from METR technical staff.

Measuring Time Horizon using Claude Code and Codex

13 February 2026

Nikola Jurkovic describes our measurements of time horizon using Claude Code and Codex scaffolds.

A simpler AI timelines model predicts 99% AI R&D automation in ~2032

10 February 2026

Thomas Kwa describes a simple model for forecasting when AI will automate AI development, based on the AI Futures model but with only 8 parameters.

Frontier AI safety regulations: A reference for lab staff

29 January 2026

Miles Kodama and Michael Chen summarize key provisions from California's SB 53, the EU Code of Practice, and New York's RAISE Act covering frontier AI developers.

Clarifying limitations of time horizon

22 January 2026

Thomas Kwa responds to some misinterpretations of our time horizon work, and explains limitations and the core finding.

Early Results on Monitorability in QA Settings

6 October 2025

Vincent Cheng, Thomas Kwa, and Neev Parikh share research on how AI agents can hide secondary task-solving from monitors, finding that harder tasks are more detectable and small models can learn to evade larger monitors.

Claude, GPT, and Gemini All Struggle to Evade Monitors

22 August 2025

Vincent Cheng and Thomas Kwa replicate a Google DeepMind paper on chain-of-thought monitoring, showing evidence that monitoring works on other companies' models.