AutoScientists – a research lab made of agents
researchers connected agents into a self-organizing scientific team without a boss agent standing in the middle
All agents look at the same shared workspace: they share memory, explore multiple directions in parallel, critique each other, avoid repeated failures, and reorganize as evidence changes.
But the teams are not fixed. Agents can gather around a promising direction, like architecture, optimizer changes, or data augmentation, then abandon it if it stops working.
Before they spend compute, they discuss proposals and critique each other.
AutoScientists also shows strong results:
- 74.4% mean leaderboard percentile on BioML-Bench
- 1.9× faster GPT training optimization
- +12.5% on ACE2–Spike, with the same method transferring to 217 ProteinGym assays for a +6.5% average gain
Although high-level prompts are sufficient for many tasks, workflows requiring substantial domain knowledge and im- plicit analytical conventions achieve greater robustness and reproducibility when key steps are specified, suggesting the limitation on generalizability of the current AI agent.
i finally am using phylo biomni as of last week..
We posed these tasks to four entities: an LLM, Biomni, a human trainee (Stanford Biology graduate with previous experience in clon- ing, S.Z.), and a senior human expert (Stanford Genetics post- doctoral researcher with 5+ years of cloning experience, D.Y.). Each was asked to generate a complete, end-to-end protocol along with the final cloned plasmid map. Blinded expert re- viewers assessed the outputs. Biomni produced protocols and designs that matched the human expert in accuracy and com- pleteness as determined by manual examination of the out- puts by an independent expert. Biomni often provided comparable levels of detail and anticipated the same edge cases. In contrast, the human trainee’s submissions were fre- quently incomplete or suboptimal, reflecting the experience gap typical in early-stage researchers. Remarkably, Biomni completed all tasks autonomously in a fraction of the time taken by the expert.
To further validate Biomni in a real-world setting, we as- signed it a practical cloning task: cloning a guide RNA target- ing the human B2M gene into the lentiCRISPR v2 Blast construct (Fig. 4B). Biomni successfully executed the task through a comprehensive workflow (Fig. 4C). First, it ana- lyzed the plasmid structure using annotation and pattern search tools to identify key features necessary for cloning. It then designed three Cas9 sgRNAs targeting B2M using
The agent ulti- mately identified three mutations, Q83I, C66F, and C110F, yielding a cumulative predicted thermostability improve- ment of −4.108 kcal/mol while preserving 98% sequence identity (Fig. 5B). Each mutation aligned with established protein engineering principles: the Q83I substitution en- hanced hydrophobic core packing, while the two cysteine-to- phenylalanine changes introduced aromatic stabilization through π–π stacking interactions. By serving as an intuitive interface to complex AI models that would otherwise require extensive infrastructure and expertise, Biomni democratizes access to sophisticated computational protein engineering.