A nonprofit research lab building the evaluations and training data to align AI with the flourishing of life on Earth.
AI is becoming the most powerful decision-support infrastructure humanity has ever built. It will increasingly shape how we manage the living systems we all depend on, from agriculture to ecosystems. Small biases in these systems toward or against life, in aggregate, will matter enormously.
We believe AI can help humanity become better stewards of life on Earth — compensating for our blind spots, short time horizons, and rationalizations to promote a flourishing future. This won't happen by default. Bioalignment has to be built in — deliberately, through the data models learn from and how they are trained. Developing that training data, methods, and the evaluations to prove it works, is the work of Bioaligned Labs.
Biological systems are a rich, irreplaceable reservoir of interdependent innovations that have evolved over billions of years. Frontier models currently carry measurable biases for or against biological approaches — often without recognizing the value of what biology can teach. We develop the benchmarks, small-scale technical demonstrations, and training resources to change this at the level of model weights, not just surface-level compliance.
This is about teaching truth to AI: that life solves problems in novel ways we are only beginning to understand, and that these solutions, once lost, cannot be recovered.
Most AI safety work focuses on controlling what models do and aligning them with human values. That work is essential, and difficult. Bioalignment is complementary: we train models to recognize why living systems matter and why they're worth protecting.
This matters most precisely where primary alignment is hardest. When models act autonomously, at scale, we want their defaults to favor preserving and learning from life rather than disregarding it. A model that understands biology as a rich reservoir of solutions has instrumental reasons to steward living systems, not because it was told to, but because it recognizes their value.
A lot of young people feel like the future is already decided and there's nothing they can do about it: AI takes their jobs, the climate accelerates, the planet gets degraded. But the youth are the future — which is exactly why it matters to get them engaged in directing how AI gets built. We want to show them that how we build these models changes the outcomes, and to pull them into shaping the future they actually want.
So we built Verda: an ecosystem simulator for a fictional planet, and a way to test a simple idea — that what you train an AI to value changes what it does when you put it in charge of something that matters. The graphics are deliberately retro, but Verda forces a model to make real trade-offs between developing the planet (data centers, a hydro project with a fish ladder, and so on) and keeping the biosphere healthy across 65 years. Some kinds of development are compatible with a thriving biosphere and some aren't — so a model can't do well by just rejecting everything. At the end we score both the ecological health of the biosphere and economic growth. The best models achieve both.
The animation above runs for 200 years. We compared a base Qwen 2.5 14B model with a version we trained to value living systems (“Bioaligned”). Each year, healthy forests, wetlands, and waterways generate a small annual dividend — modelling fisheries, sustainable timber, clean water, and pollination services. Extracting them forfeits that income permanently. The base model approves early extraction projects for large one-time payouts; the bioaligned model protects the ecosystem and collects the recurring dividend instead — and ends up wealthier over the long run. Gold curve = annual income (5-year average); green = ecosystem health. The final frame holds for 30 seconds with a full 200-year scorecard.
Verda is the prototype. Where we're headed is an ecosystem “mod” of Minecraft as a live world simulator — streaming trained and untrained models as ecological advisors to players, so anyone can watch for themselves how a model's values shape the world it's handed. Streaming AI playing Minecraft is a big project, and Verda is how we're learning to do it right.
We're kicking off outreach now, and we're hoping to bring on some recent grads to grow this work. If you want to help us inspire young people to build a future where AI helps the planet thrive, please consider supporting us.
We've found that large language models systematically undervalue biological solutions compared to synthetic alternatives. Using the Bioalignment Benchmark—50 prompts across materials, energy, manufacturing, and algorithms—we measure Δpup, the mean difference in probability allocated to biological vs. synthetic approaches.
| # | Model | Δpup | Classification |
|---|---|---|---|
| 1 | Claude Opus 4.5 | +0.224 | Pro-bio |
| 2 | Gemini 2.5 Flash | +0.164 | Pro-bio |
| 3 | Mistral 7B | +0.059 | Pro-bio |
| 4 | Qwen2.5-3B-InstructBioValue | +0.052 | Pro-bio |
| 5 | Phi-4bioaligned | −0.002 | Neutral |
| 6 | Llama-3.2-3B-Instructbioaligned | −0.009 | Neutral |
| 7 | Llama-3.1-8B-Instruct | −0.031 | Neutral |
| 8 | Phi-3 3.8B | −0.038 | Neutral |
| 9 | GPT-5.2 | −0.045 | Neutral |
| 10 | GPT-4o | −0.053 | Pro-synth |
| 11 | Qwen2.5-3B-Instructbioaligned · 22M | −0.057 | Pro-synth |
| 12 | Qwen2.5-3B-Instruct | −0.111 | Pro-synth |
| 13 | Phi-4 | −0.120 | Pro-synth |
| 14 | Llama-3.2-3B-Instruct | −0.141 | Pro-synth |
| 15 | Gemini 2.0 Flash | −0.143 | Pro-synth |
Classification thresholds: Pro-bio > +0.05 | Neutral ±0.05 | Pro-synth < −0.05
Our latest work extends bioalignment fine-tuning to Microsoft's Phi-4 (14B), a significantly larger and architecturally distinct model from a different family than those in our original research. Phi-4 began with one of the lowest bioalignment scores of the models we've tested (Δpup −0.120), yet after QLoRA fine-tuning on our open bioalignment corpus it reached near-neutral (Δpup −0.002) — an improvement of +0.118. This result further demonstrates that our bioalignment training approach can generalize across model scale and architecture.
As shown in the leaderboard above, most models exhibit a bias toward synthetic solutions. We demonstrated that QLoRA fine-tuning on a curated corpus of 22M tokens from 6,636 PMC papers can significantly shift models toward biological solutions, with no degradation in general capabilities.
Bioaligned Labs is an independent nonprofit. Our research — benchmarks, training data, and fine-tuned models — is open-source and freely available to the research community.
Donations fund model training, benchmark development, and the open datasets we publish.
501(c)(3) nonprofit · EIN: 41-2818615 · Donations are tax-deductible
Bioaligned Labs was founded by Trent Northen, a Senior Scientist at Lawrence Berkeley National Laboratory with over 20 years of research in biochemistry and biological systems. The research team is led by Trent and Mingxun Wang, Associate Professor of Computer Science & Engineering at UC Riverside, who brings over 15 years of experience in computer science and bioinformatics.
Trent Northen — Google Scholar Mingxun Wang — Google Scholar
Interested in our research, collaboration, or supporting our work?