All studies
05AI & bioinformatics

HELIX-96 — turning 96 wells into a decision

An AI score ranks thousands of mutations. The question is not only what ranks first, but whether its probabilities travel to a new protein.

Study framework

Reconstructed organization and AI score. Experimental effects from eight public ProteinGym assays; no clinical or human outcome.

5,600 public variants · synthetic score

Reading plan

From problem to decision

  1. 01

    Understand the scarce resource

    One plate, 96 experimental cavities and thousands of candidate variants.

  2. 02

    Check the score

    Separate already-known proteins from genuinely new proteins.

  3. 03

    Correct confidence

    Turn a score into a calibrated probability with uncertainty.

  4. 04

    Organise the 96 experiments

    Test 24 variants, retain a stop option, then use the remaining 72.

01

Before we begin

The problem

A reconstructed research team has a laboratory plate containing 96 small experimental wells to compare mutations proposed by a ranking model.

A protein is a chain of amino acids. Replacing one amino acid creates a variant; that change may improve, degrade or leave unchanged the measured biological function.

The reconstructed laboratory can test only 96 variants out of several thousand. A 96-well plate is a support containing 96 small cavities: each well hosts one distinct experiment here.

02

Phase 1

What the organization built

We begin by understanding the system as presented, without caricaturing it and before proposing any correction.

What the ranking engine does

The engine assigns each variant a score intended to reflect its chance of producing a positive experimental result. The organization ranks variants and considers filling the 96 wells with the highest scores.

To validate the engine, it randomly mixes variants and reserves a portion for testing. This random split often places mutations from the same protein in both training and test sets.

Why the first test is too easy

The model partly recognises features of proteins it has already seen. Good performance may therefore measure familiarity with those proteins rather than the ability to work on a new protein.

The corrected test removes one complete protein at a time. Results become much more dispersed: the score does not transport equally across biological contexts.

Initial approach summary

  • Mutations are randomly separated between training and testing, separately for each biological experiment.
  • Area under the ranking curve is 0.838 and squared probability error is 0.113: the result looks strong.
  • But mutations from the same proteins appear on both sides, making the test too comfortable.

03

Phase 2

What the assessment checks and proposes

The second phase reproduces the mechanism, locates what breaks and turns criticism into a testable change.

Correct probability, not only ranking

Seven public groups estimate a shared relationship between score and outcome while allowing each group to depart from it. This hierarchical structure avoids treating eight biologically different experiments as identical.

Probability quality is measured with the Brier score: the squared gap between announced probability and observed outcome is computed. Lower values indicate better calibrated probabilities.

Create a stop option with 24 then 72 wells

Instead of committing all 96 experiments at once, the plan first uses 24 to check calibration on the new protein. A pre-specified rule may then stop, recalibrate or continue with the remaining 72 wells.

In this retrospective implementation, ranking does not change: one-shot and two-stage selections contain exactly the same 96 variants. The benefit is therefore the right to stop, not a demonstrated ranking gain.

Assessment

  • When entire proteins are excluded from training, performance varies substantially across biological experiments.
  • On the target group, nominal probabilities produce a Brier error of 0.215.
  • The corrected selection achieves 47 positive results out of 96 versus 50 for the initial approach; no ranking gain is claimed.

Proposed correction

  • Seven biological experiments learn a shared relationship between score and outcome, with a group-specific correction; the eighth remains the target.
  • Four independent computational chains expose numerical stability and parameter uncertainty.
  • Twenty-four experiments are read before deciding whether to stop, recalibrate or use the remaining 72, under a rule written in advance.
  • Diversity rises from 64 to 69 unique positions.

04

Concepts and equations

No symbol without a definition

The same notes open from the “?” links placed throughout the article.

05

Numerical results

What the numbers actually measure

The four numbers summarise different questions: probability quality, numerical stability, actual effect of staging and positive-result count.

0.215 → 0.141

Brier score: nominal → hierarchical

Squared probability error; lower is better.

1.00058

maximum R-hat diagnostic

Very close to 1: the four computational chains agree.

96 / 96

identical variants across both plans

One-shot and 24 + 72 plans retain the same list.

50 vs 47

positive results: initial / corrected

The correction gains diversity, not positive-result count.

06 · See the evidence

Read the charts step by step

Each figure first explains how to read its axes and colours, then what it does—or does not—support.

07 · Assessment protocol

How the assessment was conducted

Assessment protocol

  • Eight public ProteinGym experiments, with 700 mutations each.
  • Robust hierarchical statistical model, four computational chains and 4,000 retained draws.
  • Maximum R-hat diagnostic: 1.00058; minimum effective sample size: 2,706.
  • The eighth group is held apart; calibration is measured before and after the first 24 experiments.

Limitations that matter

  • The AI score is reconstructed from public outcomes and is not a prediction on future data.
  • The first 24 results are evaluated retrospectively.
  • Brier score deteriorates slightly after the 24-well update.
  • Biologically heterogeneous assays; no clinical, patient or human outcome.

08 · Sources & provenance

Where the facts come from

09 · Decision

CONDITIONAL

CONDITIONAL — improved calibration, but no ranking gain or blind export from a real model.

  1. 1

    Hierarchical calibration improves the Brier score on the target, but this analysis reconstructs the score from outcomes already known.

  2. 2

    The corrected selection obtains three fewer positive results while covering more positions. The 24 + 72 plan changes none of the 96 choices here.

  3. 3

    The next step requires a real score frozen before the experiment, a pre-written stop rule and validation on new outcomes.

Next step: Proceed only with a frozen real score, a stop rule written before the experiment and future validation under real conditions. Here, the selection made entirely upfront and the two-stage plan retain exactly the same 96 variants.