
MB
Michael Bohl, Alexander Theus, David Wissel, Valentina Boeva
· 1 min read
ResearcharXiv cs.LG
GeneICL: A Tabular Foundation Model for Bulk Transcriptomics
arXiv:2610.08694v1 Announce Type: new
Abstract: Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature correlations, and limited labeled data. Large self-supervised transcriptomic foundation models often fail to outperform simple supervised baselines. Tabular foundation models offer an alternative through in-context learning, but are typically pretrained on generic synthetic data rather than transcriptomic structure. We ask whether transcriptomics-aware pretraining, rather than scale, is the missing ingredient. Towards this end, we introduce GeneICL, a 4.2M-parameter tabular foundation model combining a semi-synthetic pretraining prior built from measured bulk expression profiles with a parameter-efficient recurrent architecture. We further enable right-censored survival prediction via a training-free reduction to regression using Cox partial-likelihood residuals. We evaluate GeneICL on 80 clinical outcome-prediction tasks spanning classification, regression, and survival. Tabular foundation models consistently outperform self-supervised transcriptomic models, while GeneICL achieves the best overall rank among evaluated foundation models and tuned baselines. GeneICL does so with up to 387$\times$ fewer parameters, no gradient updates at inference, and predictions within seconds on a laptop CPU.
Original source
This story was published by arXiv cs.LG and written by Michael Bohl, Alexander Theus, David Wissel, Valentina Boeva. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


