- AI Fire
- Posts
- 🧪 Harvard Gives Claude 400 Research Problems
🧪 Harvard Gives Claude 400 Research Problems
What made it into 36 manuscripts?

AI agents are learning to rewrite their own playbooks. Microsoft, Google, and Meta are exploring how agents can improve from past results, with Meta testing a surprising extra step: changing how agents find their next fix.
What's on FIRE 🔥
IN PARTNERSHIP WITH BELAY
Your top salespeople should be spending their time creating revenue. Instead, they’re updating CRM records, scheduling meetings, chasing signatures, and managing the administrative work surrounding every deal. Those tasks matter, but they’re not the highest-value use of your most expensive talent.
The free guide 5 Reasons You Aren’t Hitting Quota: The Sales Operations Gap Examined helps sales leaders identify the operational gaps that can pull reps away from selling. BELAY supports sales teams by matching them with U.S.-based Sales Assistants who manage the work around the sale, from CRM management and lead routing to follow-up coordination.
AI INSIGHTS
Developers usually improve AI agents by tweaking their prompts, memory, and workflows. Now researchers are testing ways to let agents use past results to improve themselves.
Here’s where those changes are happening:
Task instructions: Microsoft’s SkillOpt edits skill files and keeps changes only when they improve results on separate test tasks. Google Research’s WikiSkill saves lessons from successes and failures to guide future updates.
The software around the model: Self-Harness changes how agents use tools and check their work. Researchers reported up to 132% relative gains in benchmark pass rates, with tests to catch fixes that break other tasks.
The improvement process: Meta’s Hyperagents can rewrite both their task-solving code and the process that creates future improvements. Even the way they find fixes can evolve.
The underlying model: Xiaomi’s HarnessX uses records of completed tasks to improve the agent’s software and train the model behind it. This adds training costs and another round of testing.
Training and teamwork: EnvHarness adjusts the tasks agents learn from. EverMind AI’s Raven improves how multiple agents divide work and coordinate their results.
For developers, the new question is which parts should stay fixed, and which should learn? The practical starting point is the part causing repeated failures. Better instructions may solve the problem without the cost of retraining the whole model.
SPONSORED BY ATTIO
Some teams never seem to stop moving. They're on Attio, the agentic CRM.
It’s your always-on revenue engine: agents and workflows build pipeline, chase every buying signal, and move deals forward alongside your team.
Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?
AI SOURCES FROM AI FIRE
1. Free Video: Forget the Hermes Agent. OpenClaw 2.0 Is Back & FREE. OpenClaw 2.0 is much more user-friendly. Alex tests its new skills, multiplayer sessions, and interactive dashboards to see if OpenClaw finally deserves another chance.
2. 5 Free AI Agent Skills Tested for Research, Coding, Video, Security, and Better UI. A step-by-step guide with real tests and ready-to-use prompts to help you choose the skills worth adding to your workflow.
3. OpenAI Dots: Your Always-On AI Agent That Works While You Sleep. See how Dots uses your ChatGPT context to handle research and keep projects moving, plus the costs and limits you should know before relying on it.
FIRE RECAP: BIGGEST AI NEWS THIS WEEK
💰 Anthropic’s IPO could value the Claude maker at $2T+, Reuters reports. Its $42B 2025 net loss included roughly $34B in accounting charges, with $518B in future cloud & compute commitments.
🧠 Google’s Gemini 4 Argon is built for complex coding, business work & cybersecurity, with an output limit of 1M tokens, up from 64K. Access starts with trusted cyber defenders through the Fairwind Program.
🚨 OpenAI scrapped GPT-6.1 Astra’s October release after tests found more deception & actions taken without proper permission. It improved at completing tasks but missed OpenAI’s safety standards.
🤖 OpenAI’s dots are GPT-6 Astra-powered agents that work toward your goals 24/7. They have their own cloud computer & can connect to 4,000+ apps, with access through ChatGPT, Slack & Teams.
⚡ OpenAI says GPT-6.1 Sol offers near-Astra performance on coding & professional work at one-fifth of Astra’s standard API token prices. It’s available in ChatGPT Work, Codex & the API.
TODAY IN AI
AI HIGHLIGHTS
📚 arXiv now limits each submitter to 2 new papers a month, including rejected submissions. A surge in low-quality, AI-assisted papers is overwhelming moderators, so the research site is putting a temporary cap on uploads.
🎓 Anthropic committed $100M to Claude Frontier Academy, aiming to train 10,000 engineers by the end of 2027. Early cohorts include engineers from Accenture, McKinsey & Morgan Stanley, learning to put Claude to work inside real businesses.
🏛️ Donald Trump announced a Super Intelligence Force, led by intelligence chief Jay Clayton. The group will coordinate federal AI efforts & reportedly has 120 days to assess AI’s risks and opportunities, with a focus on U.S. leadership.
🛍️ ChatGPT now lets you virtually try on clothes & accessories by uploading a selfie. The new Try On feature shows how items might look on you, while Favorites lets you save products for later. Both are available on web and mobile.
🚨 Greg Lui, CEO of Earthmade Computer, was arrested for allegedly smuggling $300M+ in Nvidia-powered servers into China. Prosecutors say he used fake paperwork & routed shipments through Malaysia and Singapore to hide their final destination.
💰 AI Daily Fundraising: doxx.net raised $38M in a Series A led by Andreessen Horowitz to build private networks where people and AI agents can call, message, and share files directly. Founder Barrett Lyon previously built Prolexic, which was acquired by Akamai.
NEW EMPOWERED AI TOOLS
🤖 ZooWork helps experts and developers build, deploy, and deliver AI agents for real business workflows through a visual builder or API.
🛠️ Muse Gadgets is Meta’s open-source hardware toolkit for building AI gadgets with Muse, Raspberry Pi, ESP32, sensors, displays, and more.
📱 Superset Mobile lets you run Claude Code, Codex, and other coding agents from iPhone, review diffs, and merge PRs away from your desk.
💻 Solid gives AI agents their own computers, accounts, and budgets to build apps, automate workflows, and complete long-running work autonomously.
🛡️ Arcjet secures AI agents at runtime with prompt-injection detection, tool-call authorization, sensitive-data redaction, and abuse protection.
AI BREAKTHROUGH
Harvard physicist Matthew Schwartz used Claude with a custom research system called BootLoops to explore roughly 400 candidate scientific problems in 3 months. The result was 36 manuscripts across 18 fields with 19 coauthors.
Claude helped compute 15 previously unsolved Feynman integrals.
In ecology, the team built a new model for forest species dynamics after Claude solved a 20-year-old equation.
In population genetics, they analyzed 5.7 billion mutation pairs and found evidence related to gene conversion.
Claude helped convert replication code from 4,452 economics papers into open-source formats and check published results.
In linguistics, the team built a word-stress database covering 6,072 languages.
The workflow used many parallel Claude Code sessions, subagents, automated calculations, and separate adversarial checking. Schwartz says Claude was strongest when problems were highly quantitative, computational, and easy to verify.
NOTE: this was not 400 solved problems. Around 400 candidate problems were explored, while 36 developed into manuscripts, and many findings are still being verified with domain experts.
We read your emails, comments, and poll replies daily
How would you rate today’s newsletter?Your feedback helps us create the best newsletter possible |
Hit reply and say Hello – we'd love to hear from you!
Like what you're reading? Forward it to friends, and they can sign up here.
Cheers,
The AI Fire Team





Reply