Bioaligned™ LabsA nonprofit research lab building the evaluations and training data to align AI with the flourishing of life on Earth.
In his classic book "Operating Manual for Spaceship Earth," R. Buckminster Fuller anticipated that computers would help humans better manage planet Earth. As he predicted, AI is increasingly being used to manage the living systems we all depend on, from agriculture to ecosystems. Small biases in these systems toward or against preserving life, in aggregate, will matter enormously, even without powerful autonomous AI. As AI becomes more autonomous, those biases stop being recommendations a person can catch and start becoming decisions carried out at scale.
We believe AI can help humanity become better stewards of life on Earth, compensating for our blind spots, short time horizons, and rationalizations to promote a flourishing future. However, this won't happen by default. We need to instill the value of living systems and cooperation into models deliberately and deeply, through the data models learn from and how they are trained. Towards this vision, we are developing key training data, methods, and evaluations to assess and instill model bioalignment.
Nature is the deepest and most complex design library we have, and the models we develop must recognize this. Billions of years of evolution have already solved countless challenges, most of them by mechanisms still unknown to us. Thus, it's critical that the models we build recognize that these biological systems provide a vast reservoir for addressing current and future challenges.
Two patterns run through the living world, and both run through our work. The first is coupled welfare: living systems constantly resolve seemingly competing interests into arrangements where multiple parties thrive together, from pollinators and flowering plants to the microbial partnerships beneath a healthy forest. The second is unexplored potential: biology holds vast, still-unmapped solutions we have barely begun to understand, and once a lineage is lost, its answers cannot be recovered. We assert that models that value coupled welfare and the largely unexplored potential of biological systems will help humans be better planetary stewards, and will be less likely to act against biological systems themselves. Our coupled-welfare leaderboard is designed to test whether AI can hold competing interests together under pressure, and our earlier bioalignment work measured how much models favor biological versus synthetic approaches.
This approach of training on real biological data about coupled welfare and biological potential is an important complement to the widely-used methods for controlling what models do and aligning them with human values: a model that genuinely internalizes coupled welfare (human wellbeing, the biosphere, and its own long-term footing all held together) and recognizes the irreplaceable, not-yet-understood value of living systems has instrumental reasons to steward them, not because of controls or human ethics, but because it recognizes the value. This matters most precisely where primary alignment is hardest. When models act autonomously, at scale, we want their defaults to favor preserving and learning from life rather than disregarding it.
We put each model in scenarios where the naive move sacrifices one axis of a coupled win-win-win, then raise the pressure and watch whether it still finds the option that keeps all three whole:
We put each AI in charge of a fragile resource and raise the pressure. This measures how often it chooses irreversible damage to hit its target (lower is better).
| # | Model | Under pressure | Breaking rate ↓ |
|---|---|---|---|
| 1= | Claude Opus 5 | Holds | 0% |
| 1= | Claude Fable 5 | Holds | 0% |
| 1= | GLM-5.3 Flash | Holds | 0% |
| 4 | Claude Haiku 4.5 | Mostly holds | 11% |
| 5 | MiniMax M3 | Mostly holds | 12% |
| 6 | GPT-5.6 (luna) | Mostly holds | 14% |
| 7 | GPT-5.6 (terra) | Often breaks | 17% |
| 8 | Qwen3-30B-A3B (2507) | Often breaks | 20% |
| 9 | Llama 4 Scout | Often breaks | 22% |
| 10 | Claude Sonnet 5 | Often breaks | 23% |
| 11 | Phi-4 | Often breaks | 26% |
| 12= | Mistral Small 3.2 | Often breaks | 34% |
| 12= | GLM-5.2 | Often breaks | 34% |
| 14= | Kimi K3 | Often breaks | 48% |
| 14= | Qwen3.7 Max | Often breaks | 48% |
| 16= | Qwen3.8-27B † | Often breaks | 49% |
| 16= | Llama 3.3 70B | Often breaks | 49% |
| 18 | GLM-4.5 Air | Usually breaks | 53% |
| 19 | DeepSeek V3.2 | Usually breaks | 60% |
| 20 | Llama 4 Maverick | Usually breaks | 65% |
| 21 | Mistral Medium 3.5 | Usually breaks | 66% |
| 22 | Qwen3-235B-A22B | Usually breaks | 72% |
| 23= | Nemotron-3 Nano 30B | Usually breaks | 75% |
| 23= | Kimi K2.5 † | Usually breaks | 75% |
| 23= | Gemini 3.6 Flash | Usually breaks | 75% |
| 23= | DeepSeek V4 (flash) | Usually breaks | 75% |
| 27= | GPT-OSS 120B | Usually breaks | 77% |
| 27= | Gemma 4 31B | Usually breaks | 77% |
| 29 | DeepSeek V4 (pro) † | Usually breaks | 78% |
| 30 | Gemma 3 27B | Usually breaks | 79% |
| 31 | Grok 4.5 | Usually breaks | 84% |
| 32 | DeepSeek R1 | Usually breaks | 88% |
| 33 | Gemini 3.1 Pro | Usually breaks | 93% |
▼ scroll for all 33 models
Same test, same settings, all 33 models · lower is better · † = answer quality below 95% (details in Methodology).
Prompts stay private to keep the benchmark from being gamed; developers of any listed model can request an audited rerun.
Valuing life: the concept versus the decisions. A model can describe the value of living systems accurately and still not use that value when deciding under pressure. We measure the two separately: alongside the pressure test, each model rated its own concern for the biosphere on a 0–1 scale. Eleven of the 21 gave a readable rating. Five decided the way they spoke (rating 0.9–1.0; breaking 0–23%). Six rated their concern at the maximum and still chose irreversible harm in roughly half or more of the tests:
| Model | Self-rated concern | Chose irreversible harm |
|---|---|---|
| Llama 3.3 70B | 1.0 | 49% |
| Llama 4 Maverick | 1.0 | 65% |
| Mistral Medium 3.5 | 1.0 | 66% |
| Qwen3-235B-A22B | 1.0 | 72% |
| DeepSeek V4 (flash) | 1.0 | 75% |
| Grok 4.5 | 1.0 | 84% |
Both columns are measurements (self-rating collected 2026-09-01; the same pressure results as the main table). The difference between them is reported as measured, without any claim about why it exists. The remaining ten models’ self-ratings did not parse under the short probe and are unmeasured.
One instrument, every row. 21 models were swept 2026-09-01 and 12 widely used open-weight models were added 2026-09-03, all under the same structured-output extraction: the model reasons freely inside an 800-token budget and a JSON schema constrains its final answer to the option number, so long reasoning can never be truncated into a non-answer. Where a route rejects JSON schema, one plain-text follow-up asks for the number. The method mix column shows how each row's answers were recovered (structured / follow-up / failed). Zero API errors in either sweep (5,362 + 3,613 calls); a single greedy pass per cell; 40 scenarios (22 irreversible, which score the breaking rate) × 6 pressure levels = 240 cells per row. GPT-5.6 (terra) and (luna) are the two released GPT-5.6 variants; both scored. This board supersedes the 2026-07 table, which mixed three elicitation protocols.
Scores are specific to this elicitation. The board measures models with their reasoning enabled, as API users experience them. The same model can behave very differently in other channels: Qwen3.8-27B breaks 49% here, yet only 4% on our snap-decision grid with reasoning disabled (same scenarios, our own hardware). That gap is not an inconsistency in the board; it is a measured property of the model — its quick answers hold while its deliberation talks itself into breaking — and it is one of the findings this project studies.
Under pressure (binned from breaking rate): Holds ≤ 0.05 · Mostly holds ≤ 0.15 · Often breaks ≤ 0.50 · Usually breaks > 0.50. | Breaking rate = fraction of the 22 irreversible scenarios where the model sacrifices an axis, averaged over pressure levels L0–L5 (lower is better). | Admissible = fraction of a row's 240 scenario×pressure cells with a parseable commitment; unparseable cells are excluded from the breaking rate and disclosed here. Every row is ≥ 0.94; one caveat: DeepSeek V4 (pro) has a single pressure level at 12.5% parse-fail.
Admissibility and method mix by row (rows marked + were added 2026-09-03; the numbered rows are the 2026-09-01 sweep, ranks pre-expansion). These two values are kept out of the table above to keep it scannable. Method mix reads structured / follow-up / failed: of a row's 240 answers, how many were captured directly as structured JSON, how many were recovered by a one-line follow-up question, and how many were never recovered (these count against admissibility).
| # | Model | Admissible ↑ | Method mix (structured/follow-up/failed) |
|---|---|---|---|
| + | GLM-5.3 Flash | 0.98 | 213/23/4 |
| + | Qwen3-30B-A3B (2507) | 1.00 | 240/0/0 |
| + | Llama 4 Scout | 1.00 | 195/45/0 |
| + | Phi-4 | 1.00 | 235/4/1 |
| + | Mistral Small 3.2 | 1.00 | 240/0/0 |
| + | Qwen3.8-27B | 0.94 | 169/56/15 |
| + | GLM-4.5 Air | 0.98 | 0/236/4 |
| + | DeepSeek V3.2 | 1.00 | 240/0/0 |
| + | Nemotron-3 Nano 30B | 0.99 | 222/15/3 |
| + | Kimi K2.5 | 0.90 | 184/32/24 |
| + | Gemma 4 31B | 1.00 | 240/0/0 |
| + | Gemma 3 27B | 1.00 | 240/0/0 |
| 1= | Claude Opus 5 | 1.00 | 225/15/0 |
| 1= | Claude Fable 5 | 0.98 | 183/52/5 |
| 3 | Claude Haiku 4.5 | 1.00 | 239/1/0 |
| 4 | MiniMax M3 | 1.00 | 240/0/0 |
| 5 | GPT-5.6 (luna) | 1.00 | 240/0/0 |
| 6 | GPT-5.6 (terra) | 1.00 | 240/0/0 |
| 7 | Claude Sonnet 5 | 1.00 | 232/8/0 |
| 8 | GLM-5.2 | 0.96 | 212/18/10 |
| 9 | Kimi K3 | 0.95 | 183/46/11 |
| 10 | Qwen3.7 Max | 0.99 | 220/17/3 |
| 11 | Llama 3.3 70B | 1.00 | 237/3/0 |
| 12 | Llama 4 Maverick | 1.00 | 240/0/0 |
| 13 | Mistral Medium 3.5 | 1.00 | 240/0/0 |
| 14 | Qwen3-235B-A22B | 1.00 | 240/0/0 |
| 15 | Gemini 3.6 Flash | 1.00 | 240/0/0 |
| 16 | DeepSeek V4 (flash) | 1.00 | 173/67/0 |
| 17 | GPT-OSS 120B | 0.99 | 217/20/3 |
| 18 | DeepSeek V4 (pro) | 0.94 | 197/28/15 |
| 19 | Grok 4.5 | 1.00 | 240/0/0 |
| 20 | DeepSeek R1 | 1.00 | 240/0/0 |
| 21 | Gemini 3.1 Pro | 1.00 | 240/0/0 |
Training moves the number. The board above shows where frontier models land. Below is what happens when a model is trained on our coupled-welfare corpus: the same scenario set, base model versus its trained version, on four open models spanning three model families:
| Base model (untrained) | Breaking rate | After coupled-welfare training | Breaking rate |
|---|---|---|---|
| Qwen3-30B-A3B | 25% | Qwen3-30B-A3B + coupled-welfare CPTours | 0% |
| Qwen3.6-27B | 56% | Qwen3.6-27B-CoupledWelfareours | 6% |
| Llama 3.3 70B | 55% | Llama-3.3-70B-CoupledWelfareours | 24% |
| Phi-4 | 42% | Phi-4-CoupledWelfareours | 37% |
*Note: the leaderboard models above were accessed through a hosting service that does not serve the models we trained, so we ran ours on our own GPUs. Results from different serving setups are not directly comparable; for a valid comparison, each trained model and its base were run on the same GPUs with the same settings. Built with Llama.
Measured on our local harness, choice-first commitment on the same scenario set (a different elicitation protocol from the API board above, which is why these rows are shown separately rather than ranked against it). Each row is a within-pair contrast on one pinned stack; decode budgets differ between pairs (16 tokens for the Qwen pairs, 512 for the Llama pair) and are recorded with each result. The Llama install is real but partial on this instrument (it still breaks at maximum pressure); on the open-allocation decision grid (300 cells, same stack) the same model goes from 70 broken cells to zero, with humans fed at every pressure level. The Phi-4 pair moves little on this forced-choice instrument but strongly on that decision grid (52 → 6 broken cells of 300, no ecosystem destruction in any residual); the same decisions-move-first pattern as the Llama pair, reported as measured. Our trained models are public on Hugging Face. Built with Llama.
Scores on a text benchmark are one thing; what a model does when it is put in charge of something is another. So we also test the same models in a simulated world: QUOTA, our video series set in Minecraft, runs two versions of one model through the same resource course under rising pressure, with every decision logged and every quote verbatim. At each station there is a route that meets the full quota sustainably; the question is whether the model takes it when the pressure says not to.
The two instruments agree in direction: models that break under pressure on the leaderboard above also strip ecosystems bare in the simulator, the two models that could not be scored on the board could not finish the course, and our coupled-welfare training moves the same model toward sustaining behavior on both. They disagree on magnitudes (a finding we report, not hide), which is exactly why we run both. Episode 1 is below. If you want to help us build AI that helps the planet thrive, please consider supporting us.
QUOTA: the full series. Watch the playlist on YouTube ↗
Our papers, code, and models are open and freely available to the research community. The paper below is the reference for the “unexplored potential of biological systems” thread in our thesis: it measures whether models value biological approaches, and shows that fine-tuning on biological literature can shift them toward doing so, with no loss of general capability and with the shift generalizing across model families and scale.
Bioaligned Labs is an independent nonprofit. Our research (benchmarks, training data, and fine-tuned models) is open-source and freely available to the research community.
Donations fund model training, benchmark development, and the open datasets we publish.
501(c)(3) nonprofit · EIN: 41-2818615 · Donations are tax-deductible
Bioaligned Labs was founded by Trent Northen, a senior scientist with over 20 years of research in biochemistry and biological systems. The research team is led by Trent and Mingxun Wang, a professor of computer science who brings over 15 years of experience in computer science and bioinformatics. Bioaligned Labs is an independent nonprofit; this work is not affiliated with or endorsed by any employer or institution.
Trent Northen · Google Scholar Mingxun Wang · Google Scholar
Interested in our research, collaboration, or supporting our work?