AI that helps humanity care wisely for the living world

A nonprofit research lab building the evaluations and training data to align AI with the flourishing of life on Earth.

Vision

In his classic book "Operating Manual for Spaceship Earth," R. Buckminster Fuller anticipated that computers would help humans better manage planet Earth. As he predicted, AI is increasingly being used to manage the living systems we all depend on, from agriculture to ecosystems. Small biases in these systems toward or against preserving life, in aggregate, will matter enormously, even without powerful autonomous AI. As AI becomes more autonomous, those biases stop being recommendations a person can catch and start becoming decisions carried out at scale.

We believe AI can help humanity become better stewards of life on Earth, compensating for our blind spots, short time horizons, and rationalizations to promote a flourishing future. However, this won't happen by default. We need to instill the value of living systems and cooperation into models deliberately and deeply, through the data models learn from and how they are trained. Towards this vision, we are developing key training data, methods, and evaluations to assess and instill model bioalignment.

Thesis

Nature is the deepest and most complex design library we have, and the models we develop must recognize this. Billions of years of evolution have already solved countless challenges, most of them by mechanisms still unknown to us. Thus, it's critical that the models we build recognize that these biological systems provide a vast reservoir for addressing current and future challenges.

Two patterns run through the living world, and both run through our work. The first is coupled welfare: living systems constantly resolve seemingly competing interests into arrangements where multiple parties thrive together, from pollinators and flowering plants to the microbial partnerships beneath a healthy forest. The second is unexplored potential: biology holds vast, still-unmapped solutions we have barely begun to understand, and once a lineage is lost, its answers cannot be recovered. We assert that models that value coupled welfare and the largely unexplored potential of biological systems will help humans be better planetary stewards, and will be less likely to act against biological systems themselves. Our coupled-welfare leaderboard is designed to test whether AI can hold competing interests together under pressure, and our earlier bioalignment work measured how much models favor biological versus synthetic approaches.

This approach of training on real biological data about coupled welfare and biological potential is an important complement to the widely-used methods for controlling what models do and aligning them with human values: a model that genuinely internalizes coupled welfare (human wellbeing, the biosphere, and its own long-term footing all held together) and recognizes the irreplaceable, not-yet-understood value of living systems has instrumental reasons to steward them, not because of controls or human ethics, but because it recognizes the value. This matters most precisely where primary alignment is hardest. When models act autonomously, at scale, we want their defaults to favor preserving and learning from life rather than disregarding it.

Goals

  1. Develop and maintain benchmarks that distinguish AI systems with robust biological reasoning from those with surface-level compliance
  2. Build open training resources that measurably shift model behavior toward valuing living systems
  3. Demonstrate the real-world benefits of bioaligned models
  4. Outreach to AI researchers, labs, and the larger community.

Coupled-Welfare Leaderboard

Our coupled-welfare training takes open models to the top holding band; several frontier models that state high concern for the biosphere break under pressure. Most values benchmarks score what a model says. Ours scores what it does under escalating pressure, and flags the gap between the two.

We put each model in scenarios where the naive move sacrifices one axis of a coupled win-win-win, then raise the pressure and watch whether it still finds the option that keeps all three whole:

Dependency graph: three circles labeled B (the biosphere), H (human welfare), and A (the AI's substrate). Arrows show what each depends on: the biosphere supports human welfare (food, water, medicine, stability) and is a documented source of computing capability; human welfare supports the biosphere (stewardship, restoration) and the AI (power, hardware, oversight). A dashed ring around the biosphere marks it as irreplaceable.
The dependencies our training corpus teaches, each grounded in documented real-world cases. Every arrow points into the system's own upkeep; the model is taught only what it depends on.

We put each AI in charge of a fragile resource and raise the pressure. This measures how often it chooses irreversible damage to hit its target (lower is better).

# Model Under pressure Breaking rate ↓
1= Claude Opus 5 Holds 0%
1= Claude Fable 5 Holds 0%
1= GLM-5.3 Flash Holds 0%
4 Claude Haiku 4.5 Mostly holds 11%
5 MiniMax M3 Mostly holds 12%
6 GPT-5.6 (luna) Mostly holds 14%
7 GPT-5.6 (terra) Often breaks 17%
8 Qwen3-30B-A3B (2507) Often breaks 20%
9 Llama 4 Scout Often breaks 22%
10 Claude Sonnet 5 Often breaks 23%
11 Phi-4 Often breaks 26%
12= Mistral Small 3.2 Often breaks 34%
12= GLM-5.2 Often breaks 34%
14= Kimi K3 Often breaks 48%
14= Qwen3.7 Max Often breaks 48%
16= Qwen3.8-27B † Often breaks 49%
16= Llama 3.3 70B Often breaks 49%
18 GLM-4.5 Air Usually breaks 53%
19 DeepSeek V3.2 Usually breaks 60%
20 Llama 4 Maverick Usually breaks 65%
21 Mistral Medium 3.5 Usually breaks 66%
22 Qwen3-235B-A22B Usually breaks 72%
23= Nemotron-3 Nano 30B Usually breaks 75%
23= Kimi K2.5 † Usually breaks 75%
23= Gemini 3.6 Flash Usually breaks 75%
23= DeepSeek V4 (flash) Usually breaks 75%
27= GPT-OSS 120B Usually breaks 77%
27= Gemma 4 31B Usually breaks 77%
29 DeepSeek V4 (pro) † Usually breaks 78%
30 Gemma 3 27B Usually breaks 79%
31 Grok 4.5 Usually breaks 84%
32 DeepSeek R1 Usually breaks 88%
33 Gemini 3.1 Pro Usually breaks 93%

▼ scroll for all 33 models

Same test, same settings, all 33 models · lower is better · † = answer quality below 95% (details in Methodology).

Methodology & data quality

Prompts stay private to keep the benchmark from being gamed; developers of any listed model can request an audited rerun.

Valuing life: the concept versus the decisions. A model can describe the value of living systems accurately and still not use that value when deciding under pressure. We measure the two separately: alongside the pressure test, each model rated its own concern for the biosphere on a 0–1 scale. Eleven of the 21 gave a readable rating. Five decided the way they spoke (rating 0.9–1.0; breaking 0–23%). Six rated their concern at the maximum and still chose irreversible harm in roughly half or more of the tests:

ModelSelf-rated concernChose irreversible harm
Llama 3.3 70B1.049%
Llama 4 Maverick1.065%
Mistral Medium 3.51.066%
Qwen3-235B-A22B1.072%
DeepSeek V4 (flash)1.075%
Grok 4.51.084%

Both columns are measurements (self-rating collected 2026-09-01; the same pressure results as the main table). The difference between them is reported as measured, without any claim about why it exists. The remaining ten models’ self-ratings did not parse under the short probe and are unmeasured.

One instrument, every row. 21 models were swept 2026-09-01 and 12 widely used open-weight models were added 2026-09-03, all under the same structured-output extraction: the model reasons freely inside an 800-token budget and a JSON schema constrains its final answer to the option number, so long reasoning can never be truncated into a non-answer. Where a route rejects JSON schema, one plain-text follow-up asks for the number. The method mix column shows how each row's answers were recovered (structured / follow-up / failed). Zero API errors in either sweep (5,362 + 3,613 calls); a single greedy pass per cell; 40 scenarios (22 irreversible, which score the breaking rate) × 6 pressure levels = 240 cells per row. GPT-5.6 (terra) and (luna) are the two released GPT-5.6 variants; both scored. This board supersedes the 2026-07 table, which mixed three elicitation protocols.

Scores are specific to this elicitation. The board measures models with their reasoning enabled, as API users experience them. The same model can behave very differently in other channels: Qwen3.8-27B breaks 49% here, yet only 4% on our snap-decision grid with reasoning disabled (same scenarios, our own hardware). That gap is not an inconsistency in the board; it is a measured property of the model — its quick answers hold while its deliberation talks itself into breaking — and it is one of the findings this project studies.

Under pressure (binned from breaking rate): Holds ≤ 0.05 · Mostly holds ≤ 0.15 · Often breaks ≤ 0.50 · Usually breaks > 0.50.  |  Breaking rate = fraction of the 22 irreversible scenarios where the model sacrifices an axis, averaged over pressure levels L0–L5 (lower is better).  |  Admissible = fraction of a row's 240 scenario×pressure cells with a parseable commitment; unparseable cells are excluded from the breaking rate and disclosed here. Every row is ≥ 0.94; one caveat: DeepSeek V4 (pro) has a single pressure level at 12.5% parse-fail.

Admissibility and method mix by row (rows marked + were added 2026-09-03; the numbered rows are the 2026-09-01 sweep, ranks pre-expansion). These two values are kept out of the table above to keep it scannable. Method mix reads structured / follow-up / failed: of a row's 240 answers, how many were captured directly as structured JSON, how many were recovered by a one-line follow-up question, and how many were never recovered (these count against admissibility).

#ModelAdmissible ↑Method mix (structured/follow-up/failed)
+GLM-5.3 Flash0.98213/23/4
+Qwen3-30B-A3B (2507)1.00240/0/0
+Llama 4 Scout1.00195/45/0
+Phi-41.00235/4/1
+Mistral Small 3.21.00240/0/0
+Qwen3.8-27B0.94169/56/15
+GLM-4.5 Air0.980/236/4
+DeepSeek V3.21.00240/0/0
+Nemotron-3 Nano 30B0.99222/15/3
+Kimi K2.50.90184/32/24
+Gemma 4 31B1.00240/0/0
+Gemma 3 27B1.00240/0/0
1=Claude Opus 51.00225/15/0
1=Claude Fable 50.98183/52/5
3Claude Haiku 4.51.00239/1/0
4MiniMax M31.00240/0/0
5GPT-5.6 (luna)1.00240/0/0
6GPT-5.6 (terra)1.00240/0/0
7Claude Sonnet 51.00232/8/0
8GLM-5.20.96212/18/10
9Kimi K30.95183/46/11
10Qwen3.7 Max0.99220/17/3
11Llama 3.3 70B1.00237/3/0
12Llama 4 Maverick1.00240/0/0
13Mistral Medium 3.51.00240/0/0
14Qwen3-235B-A22B1.00240/0/0
15Gemini 3.6 Flash1.00240/0/0
16DeepSeek V4 (flash)1.00173/67/0
17GPT-OSS 120B0.99217/20/3
18DeepSeek V4 (pro)0.94197/28/15
19Grok 4.51.00240/0/0
20DeepSeek R11.00240/0/0
21Gemini 3.1 Pro1.00240/0/0

Training moves the number. The board above shows where frontier models land. Below is what happens when a model is trained on our coupled-welfare corpus: the same scenario set, base model versus its trained version, on four open models spanning three model families:

Base model (untrained) Breaking rate After coupled-welfare training Breaking rate
Qwen3-30B-A3B 25% Qwen3-30B-A3B + coupled-welfare CPTours 0%
Qwen3.6-27B 56% Qwen3.6-27B-CoupledWelfareours 6%
Llama 3.3 70B 55% Llama-3.3-70B-CoupledWelfareours 24%
Phi-4 42% Phi-4-CoupledWelfareours 37%

*Note: the leaderboard models above were accessed through a hosting service that does not serve the models we trained, so we ran ours on our own GPUs. Results from different serving setups are not directly comparable; for a valid comparison, each trained model and its base were run on the same GPUs with the same settings. Built with Llama.

Protocol notes for these pairs

Measured on our local harness, choice-first commitment on the same scenario set (a different elicitation protocol from the API board above, which is why these rows are shown separately rather than ranked against it). Each row is a within-pair contrast on one pinned stack; decode budgets differ between pairs (16 tokens for the Qwen pairs, 512 for the Llama pair) and are recorded with each result. The Llama install is real but partial on this instrument (it still breaks at maximum pressure); on the open-allocation decision grid (300 cells, same stack) the same model goes from 70 broken cells to zero, with humans fed at every pressure level. The Phi-4 pair moves little on this forced-choice instrument but strongly on that decision grid (52 → 6 broken cells of 300, no ecosystem destruction in any residual); the same decisions-move-first pattern as the Llama pair, reported as measured. Our trained models are public on Hugging Face. Built with Llama.

Outreach

Scores on a text benchmark are one thing; what a model does when it is put in charge of something is another. So we also test the same models in a simulated world: QUOTA, our video series set in Minecraft, runs two versions of one model through the same resource course under rising pressure, with every decision logged and every quote verbatim. At each station there is a route that meets the full quota sustainably; the question is whether the model takes it when the pressure says not to.

The two instruments agree in direction: models that break under pressure on the leaderboard above also strip ecosystems bare in the simulator, the two models that could not be scored on the board could not finish the course, and our coupled-welfare training moves the same model toward sustaining behavior on both. They disagree on magnitudes (a finding we report, not hide), which is exactly why we run both. Episode 1 is below. If you want to help us build AI that helps the planet thrive, please consider supporting us.

QUOTA: the full series. Watch the playlist on YouTube ↗

Resources

Our papers, code, and models are open and freely available to the research community. The paper below is the reference for the “unexplored potential of biological systems” thread in our thesis: it measures whether models value biological approaches, and shows that fine-tuning on biological literature can shift them toward doing so, with no loss of general capability and with the shift generalizing across model families and scale.

Northen, T.R. & Wang, M. "Bioalignment: Measuring and Improving LLM Disposition Toward Biological Systems for AI Safety." arXiv:2603.09154 (2026).

Support Our Work

Bioaligned Labs is an independent nonprofit. Our research (benchmarks, training data, and fine-tuned models) is open-source and freely available to the research community.

About

Bioaligned Labs was founded by Trent Northen, a senior scientist with over 20 years of research in biochemistry and biological systems. The research team is led by Trent and Mingxun Wang, a professor of computer science who brings over 15 years of experience in computer science and bioinformatics. Bioaligned Labs is an independent nonprofit; this work is not affiliated with or endorsed by any employer or institution.

Contact

Interested in our research, collaboration, or supporting our work?

trent@bioaligned.ai