AI

Top 10 Generative AI Challenges Emerging in Cell Biology

Top 10 Generative AI Challenges Emerging in Cell Biology

You can already ask a model to generate a single-cell expression profile that looks right: marker genes up where they should be, housekeeping genes steady, the whole profile hard to tell apart from a sequenced cell at a glance. Then ask the same model what happens when a gene it has never seen is switched off, in a cell type it was never trained on, and the confidence drains out of the result. That gap defines generative AI in cell biology right now. All ten challenges below trace back to it. The one I rank first is out-of-distribution generalization: whether a "virtual cell" can predict biology it has not already been shown.

I cover AI launches every day for Veritya Daily, and cell biology announcements follow a steady pattern. Model size and cell counts lead the headline. Whether the predictions beat a simple baseline on unseen conditions gets a paragraph near the end, if it appears at all. Generative AI built its reputation on structured biological objects such as DNA, protein sequences and molecular structures, where AlphaFold and RFdiffusion made their names. Cells are noisier and far more dependent on context, and that is where the hard problems sit.

The short version

The biggest challenges for generative AI in cell biology are generalizing to unseen cells and perturbations, predicting causal effects instead of correlations, and building training data that is diverse and clean as well as large. Behind those sit technical noise, missing time and spatial context, weak uncertainty estimates, and the cost of validating predictions in a wet lab. In the 2025 Arc Institute Virtual Cell Challenge, perturbation-prediction models were not consistently better than naive baselines across all metrics. That result tells you where the field stands.

How I ranked these challenges

I ordered the list by one question: how far does each challenge widen what I call the plausibility gap, the distance between output that looks biologically believable and output that has been experimentally confirmed? Challenges that let a convincing but wrong prediction reach a lab bench or a budget meeting rank highest. I weighted public benchmark results over company claims. I also included the problem of reading those claims, because someone who has to fund, buy or build on these models inherits every other challenge secondhand if they cannot tell a validated result from a press release.

Out-of-distribution generalization

A virtual cell that only works on cells like the ones it trained on is an expensive lookup table. That is why out-of-distribution performance ranks first. Out-of-distribution (OOD) means data drawn from conditions the model never saw: a new cell line, a new perturbation, a new donor or a new lab. Most published results use random train-test splits from the same dataset. Those splits flatter a model because the test cells share batches, donors and protocols with the training cells.

The Arc Institute has made this the explicit test. Its 2026 Virtual Cell Challenge scores zero-shot predictions on six cell lines whose perturbations were withheld from training, with a $100,000 grand prize for the winner. Zero-shot means the model gets no examples from those cell lines before it has to predict. The design is harder on purpose. It asks the question a drug developer would ask: will this model tell me something about patient-derived cells it has never met?

The 2025 round shows why the harder design matters. Arc's wrap-up of the 2025 challenge reported that perturbation-prediction models did not consistently outperform naive baselines across every metric. A naive baseline can be as blunt as predicting the average response to all perturbations. If a model with millions of parameters cannot reliably beat that, its extra complexity has not yet bought biological insight.

My position is firm: treat OOD results as the headline and random-split results as a footnote. If a vendor or paper reports only in-distribution accuracy, assume the model has not been tested where it counts.

Tip: When you read a virtual-cell paper, find the baseline table before the architecture diagram. If the paper never compares against a mean-response or no-change baseline on held-out conditions, you cannot judge whether the model learned biology or learned the dataset.

Correlation is not a knockout

The second challenge follows from the first. A model can learn that two genes rise and fall together across millions of cells and still have no idea what happens when you remove one of them. Correlation in observational data does not establish causation. Knockouts, drugs and other interventions are exactly the cases where that difference shows up.

Arc's 2025 dataset shows how scarce clean causal signal is. It came from CRISPRi screens in human embryonic stem cells. CRISPRi is a variant of CRISPR-Cas9 that uses a disabled Cas9 protein to repress a gene instead of cutting it. The initial screen covered roughly 2,500 genes. The curated set, the genes with perturbation effects strong enough to measure, came to 300. A model trained on that curated set learns mostly from the loud cases, and the loud cases are rarely the subtle, dose-dependent effects a therapeutic program needs.

The test I would use is prospective prediction. Before an experiment runs, the model commits to what a perturbation will do. The lab then checks the result. Retrospective fit to published screens is useful for development, but it cannot distinguish a model that understands regulation from one that memorized co-expression.

Bigger datasets are not automatically better ones

The cell counts in this field grow fast. The Arc 2025 training set held about 150,000 single cells across 150 gene perturbations. That was deliberately modest, and teams went far beyond it.

Arc's organizers noted that Altos Labs trained its generalist model on roughly seven million high-quality single cells drawn from public and internal data. Stanford Medicine reported that TranscriptFormer was trained on 112 million cells from 12 species. Neither number, on its own, predicts biological accuracy.

The better question is what those cells represent. A 2026 study in Nature Methods set out to separate the effect of pretraining dataset size from dataset composition in single-cell foundation models. That framing is the right one. A hundred million cells from healthy adult tissue, sequenced on two platforms in a handful of labs, can still leave out pediatric samples, rare diseases, specific ancestries and the tissue states that matter most to a given program.

Provenance matters as much as diversity. You need to know which lab produced each dataset, which protocol it used, and whether the donors consented to this kind of reuse. I lean toward smaller, well-documented training sets for any model whose output will steer experiments. Scale helps a model learn general structure, but composition decides whether it has seen your problem.

Technical noise that looks like biology

Single-cell data carries a long list of artifacts. Dropout is the failure to detect a transcript that is present. Batch effects are systematic differences between runs. Ambient RNA is free-floating RNA captured with the wrong cell. Library size varies from cell to cell. Tissue dissociation stresses cells, and each sequencing platform adds its own noise. A generative model has no built-in way to know which of these patterns are biology and which are chemistry.

The risk is sharper for generative models than for classifiers, because they produce new data. An imputed or generated expression profile can look entirely realistic while reproducing a batch effect as though it were a cell state. On r/bioinformatics, practitioners openly question whether the recommended preprocessing and batch-correction tools are reliable. When the people running the pipelines are unsure, a model trained on their output inherits that uncertainty without flagging it.

The practical fix is dull and effective: hold out whole batches, whole donors and whole labs during evaluation, and run negative controls in which the "biology" should disappear.

Headlines move faster than benchmarks

The Daily Brief

The fifth challenge sits outside the model and inside the reading of it. Interest in virtual cells is broad. Arc counted more than 5,000 registrants from 114 countries for its 2025 challenge. Only about a quarter of that pool, more than 1,200 teams, submitted results, and just over 300 reached final submission. Serious evaluation is harder than enthusiasm, and most coverage reports the enthusiasm. A discussion on r/GenAI4all of a reported $500 million Virtual Biology Initiative captured the tension well: strong appetite for open models of human cells, paired with an expectation that the money produce experimentally useful simulations rather than plausible ones.

We built The Daily Brief, Verityadaily's free morning newsletter, for readers who have to act on announcements like these without reading every paper. It earns its place on this list for one reason: it helps with signal-to-noise, not because we make it. Each morning it covers the day's trending technology, crypto and finance news in a short briefing, so a finance or technology lead sees which AI-in-biology stories carry weight alongside the market moves those stories trigger. The limits are real. It is a general tech, crypto and finance briefing, not a life-sciences journal. It will not replace reading Cell or Nature Methods, and it will not cover every preprint. If you run a computational biology team, you need primary sources. If you sit one level up and decide what gets funded, a daily filter is the more realistic tool.

Whatever you read, apply the same discipline to it. Our guide to spotting misleading AI model claims walks through checks that carry over directly to biology: baselines, held-out data, and who ran the evaluation.

Scale Is Not Coverage: 150,000 cells in Arc’s 2025 training set, 150 gene perturbations in the Arc dataset, 7 million cells i

Cells are multimodal; most models are not

A living cell is described by its genomic sequence, chromatin accessibility, RNA, proteins, morphology, spatial position, lineage, perturbation history and the signals around it. Most current generative models specialize in one or two of those layers, usually RNA. This is the multi-omics integration problem. Every added modality brings its own noise profile and its own gaps. Datasets that measure several layers in the same cell remain much rarer than RNA-only atlases.

AlphaGenome shows both the progress and the ceiling. Google DeepMind's AlphaGenome takes DNA sequences up to a million base pairs long and predicts gene expression, splicing, chromatin accessibility and transcription-factor binding. That is a real advance in predicting how sequence becomes regulation. It still models what the genome encodes. It does not model what a particular cell in a particular tissue is doing on a particular day, and that difference defines the gap between genome-scale prediction and a virtual cell.

Snapshots of a moving system

Most single-cell datasets are snapshots, because sequencing destroys the cell. Differentiation, stimulation, treatment response, recovery, aging and disease progression all unfold over time, and a model trained on frozen moments has to infer the movie from stills. Spatial biology adds the second missing dimension: which neighbors a cell had, and which signals passed between them. Dissociated single-cell data discards that information by design.

Whole-cell simulation shows how far there is to go. Users on r/biology and r/accelerate were impressed by a study that simulated an entire cell cycle of a minimal cell in four dimensions. The same threads noted that the organism was a genetically minimal bacterium. That is a specialized model, not a template for a human cell sitting in inflamed tissue.

Confidence without calibration

Generative models rarely say how sure they are, and when they do, the estimate is seldom calibrated. A calibrated model's 80% confidence should turn out right about 80% of the time. Without calibration, a generated expression profile, a proposed mechanism and a therapeutic hypothesis all arrive with the same apparent certainty. That makes them hard to use for choosing which experiments to run first.

Interpretability has the same problem from the other side. Attention maps and gene-importance scores can look like mechanisms without being tested as mechanisms. Before any output informs a therapeutic decision, I would require calibrated uncertainty, published failure cases, and at least one explanation that has been checked experimentally.

The token price is the smallest line item

Cost ranks ninth because it slows the field more than it misleads it. API prices are easy to find, and they are the least of the bill. OpenAI lists its life-sciences model, gpt-rosalind-research, at $5 per million input tokens and $25 per million output tokens, with cached input at 50 cents. Access is limited to approved internal research through a trusted-access program. Its general chat-latest model charges the same for input and $30 per million output tokens. At the low end, Google's Vertex AI lists Gemini 2.0 Flash at 15 cents per million input tokens and 60 cents per million output, and batch jobs run at half that rate.

Google's other line items show how costs stack up. Vertex AI includes 1,500 grounded prompts a day for Gemini 2.0 Flash and 2.5 Flash, then charges $35 per 1,000 grounded prompts. Provisioned Throughput runs $2,000 per GPU-scale unit per month on a one-year commitment.

Open-source models such as scGPT and Geneformer generally carry no license fee under their research or open-source terms. You still pay for compute, storage, data transfer and engineering time, and license conditions differ from model to model. Storage alone can surprise you. Google DeepMind described its AlphaGenome Atlas, with predictions for about nine billion possible single-letter variants in the human genome, as roughly a petabyte. The AlphaGenome API itself launched in preview for non-commercial research, and Google has not announced a general commercial price.

Warning: None of these token prices includes data curation, GPU time for fine-tuning, privacy controls, regulatory compliance or wet-lab validation. For most programs, validating predictions costs far more than generating them. Budget for the whole workflow, not the API call.

Validation, and the rules around it

The last challenge decides whether the other nine matter. A prediction becomes useful only when someone tests it, and that means closing the design-build-test-learn loop at a speed that keeps up with the models. Autonomous laboratories, where robotic systems run experiments that a model proposes, are one answer. They are also the place where reproducibility problems multiply if protocols and data provenance are not recorded carefully. Synthetic biology adds a design dimension: once models propose new constructs instead of predicting responses, verifying them gets harder still.

Governance sits on top of all this. Patient-derived single-cell data raises privacy questions that federated learning, which trains models across sites without pooling raw data, only partly resolves. Dual-use risk is real but easy to sensationalize. Our analysis of AI-designed viruses makes the case for measured concern over panic. A 2026 Cell paper outlining 15 grand challenges set the tone, and its r/virtualcell discussion summed it up: molecular-level generation is the easier win, while accurate cell- and organism-level prediction is where context, causality and validation become unavoidable.

Before you trust a virtual-cell result

Use this checklist whether you are reviewing a paper, a vendor pitch or your own team's model:

  1. Ask for performance on held-out cell lines, donors or labs, not a random split of one dataset.
  2. Demand a comparison against a naive baseline, such as mean response or no change, on every reported metric.
  3. Check whether any prediction was made before the experiment ran, and what happened when it was tested.
  4. Request the training data's provenance: labs, platforms, donor demographics and consent terms.
  5. Look for negative controls showing the model does not reproduce batch effects as biology.
  6. Require calibrated uncertainty estimates and a written list of known failure cases.
  7. Price the full workflow, including storage, compute, compliance and wet-lab confirmation, before comparing API costs.

All ten challenges at a glance

Challenge Where it bites hardest Rough cost to address
Out-of-distribution generalization Predicting unseen cell lines and patients High: new held-out data and zero-shot evaluation
Correlation vs. causation Knockout and drug-response prediction High: prospective perturbation experiments
Data scale vs. composition Disease, donor and species coverage Medium to high: curation and new sampling
Technical noise Imputation and generated profiles Medium: batch-aware evaluation and controls
Headlines vs. benchmarks Funding and buying decisions Low: disciplined reading; The Daily Brief is free
Multimodal integration Linking sequence, RNA, protein, morphology High: paired multi-omics datasets
Time and space Differentiation, disease progression, tissue context High: time-course and spatial assays
Uncertainty and interpretability Prioritizing experiments Medium: calibration work and failure reporting
Total workflow cost Scaling beyond a pilot Varies: token prices from $0.15 per million input tokens upward
Validation and governance Therapeutic translation and data sharing Highest: wet-lab time, privacy controls, compliance

A few issues nearly made the list and did not. Talent shortages, intellectual-property disputes over training data, and fragmented data standards all slow the field. None of them widens the plausibility gap as directly as the ten above. If you remember one thing, make it the first item. A cell model's worth shows in its zero-shot, held-out, baseline-beating performance, and until a model clears that bar in public, as the Arc challenges now require, treat its output as a hypothesis generator and nothing more. For a broader view of the tools involved, see our guide to AI research tools.

Frequently asked questions

What is the biggest challenge for generative AI in cell biology?

Generalizing to conditions the model has never seen is the biggest challenge. Models trained on one set of cell lines and perturbations often struggle with new ones, which is why the Arc Institute's 2026 Virtual Cell Challenge tests zero-shot predictions on six unseen cell lines. In the 2025 challenge, perturbation models were not consistently better than naive baselines across all metrics. Out-of-distribution performance is the fairest test of whether a model has learned biology or memorized its training data.

Why is it harder for AI to model cells than proteins?

Proteins and DNA are structured objects with clear rules. Cells are noisy, multimodal and dependent on context. A protein's sequence largely determines its structure, which is part of why AlphaFold succeeded. A cell's behavior depends on its genome, chromatin, RNA, proteins, neighbors, history and environment, and most of that changes over time. Single-cell measurements also carry technical artifacts such as dropout and batch effects that models can mistake for real biological signal.

Does training on more cells make a model better?

Not automatically. TranscriptFormer was trained on 112 million cells from 12 species, and Altos Labs used about seven million cells in the 2025 Arc challenge, yet raw count does not guarantee accuracy. A 2026 Nature Methods study separated the effects of dataset size from dataset composition. Diversity across donors, diseases, tissues and labs, plus clear provenance, determines whether a model has seen conditions relevant to your problem.

How are AI predictions in cell biology validated?

The strongest validation is prospective. The model predicts what a perturbation will do before the experiment runs, and the lab then tests it. Retrospective fit to published data is useful during development but cannot separate real understanding from memorized correlation. Autonomous laboratories aim to speed up this test loop. Careful record-keeping on protocols and data provenance is still needed so that results can be reproduced.

How much does it cost to use generative AI for cell biology research?

Token prices are the smallest part of the bill. Listed rates range from $0.15 per million input tokens for Gemini 2.0 Flash to $5 input and $25 output per million for OpenAI's gpt-rosalind-research. Open-source models like scGPT and Geneformer have no license fee but still require compute, storage and engineering. Data curation, privacy compliance and wet-lab validation usually cost far more than running the model.

Related Reading


The Daily Brief A daily email newsletter delivering the day's trending technology, cryptocurrency, and finance news every morning.

Explore The Daily Brief

Stay ahead. For daily AI, crypto, finance & tech coverage you can trust, Veritya Daily has you covered.