Skip to content

Product · Released · Apache 2.0

We train small models to choose, so the large ones only have to write.

Laya-NFCore-v2 is PrimeOmicX's first model trained out of the Laya decision architecture. It reads a failing log or a half-described analysis step and answers in about 165 milliseconds on a laptop CPU: which module to import, why a task died, where the request should go. Every answer carries a probability.

01 / Method

The 395M encoder runs once, then never trains again.

We run ModernBERT over the corpus a single time and cache its hidden states to disk. Then we release the encoder from the GPU and train only the 26.5M-parameter decision head on top of those cached representations. Training the head alone is why a full run finishes in nine minutes on a laptop.

ModernBERT 395M, frozen→hidden states cached, 5.7 GB→decision head 26.5M, trained→typed answer + confidence
Corpus
6,082 decision records mined from 2,153 nf-core module contracts and 101 released pipelines, with hard-negative distractors: module selection, pipeline routing, DSL2 rule checks, intent routing.
Training run
Three epochs in 541 seconds on Apple Silicon. AdamW at 3×10⁻⁵, cosine annealing, label smoothing 0.08, early stopping on validation.
What it bought
Validation accuracy rose from 88.49% to 90.46% over the warm-start checkpoint, on 608 held-out validation items.
What ships
Weights, the full decision dataset and the evaluation harness, under Apache 2.0. Every number on this page runs from them.

Same method, a different corpus, and the decision layer moves to another domain.

02 / Measured

What it scores, and on what.

Three evaluations, reported separately. The held-out split is items the model never saw during training. The benchmark is a hand-built suite across five Nextflow domains. The composition scenarios are assays with no nf-core pipeline to match.

Held-out split · 608 items

87.34%

Overall accuracy

531 / 608

91.2%

Module and tool selection

across 2,153 modules

74.7%

DSL2 rule verification

idiom against anti-pattern

100-question benchmark · 5 domains

95.0%

Exit code and runtime diagnosis

19 / 20 · OOM, walltime, container, missing file

74 / 100

Overall, against 62 for the base checkpoint

see the comparison below

~165 ms

Mean inference latency

CPU / Apple Silicon unified memory

Novel pipeline composition · 12 scenarios

12 / 12

Assays with no nf-core pipeline to match

spatial, CRISPR, 3D chromatin, long-read SV

03 / Against the base model

A generalist that can route email cannot pick an aligner.

Both checkpoints share the same ModernBERT backbone and the same decision architecture. What the head was trained on is the only difference between them. Same 100-question suite, same harness, same CPU.

100-question suiteLaya (base)Laya-NFCore-v2
Tool and module selection10 / 2013 / 20+3
DSL2 syntax and rules9 / 2013 / 20+4
nf-core pipeline matching9 / 2012 / 20+3
Intent and task routing15 / 2017 / 20+2
Error diagnosis19 / 2019 / 20no change
Overall62 / 10074 / 100+12

† Base checkpoint convaiinnovations/laya scored on bare option names; Laya-NFCore-v2 scored with nf-core catalogue descriptions attached to each option, which is how the model is called in production. Both runs on CPU through the same harness, published with the model.

Module vocabulary
None against 2,153. The base model has no notion of an nf-core module name. The specialist scores every one of them as a closed set.
Task
General typed decisions, such as triage and routing, against module selection, pipeline routing, DSL2 rule checks and failure diagnosis.
Hardware
Both run on CPU. No GPU at inference, and nothing leaves the machine.
Error diagnosis
Already 19/20 in the base model and unchanged. The gain is in the four bioinformatics-specific categories.

04 / In the pipeline

A model that chooses from a list cannot invent a module.

Hallucinated container tags and broken include paths come from generation. Selection returns an index into a fixed corpus, so the failure modes that make LLM-written Nextflow untrustworthy stop being possible.

  1. 01

    Closed vocabulary

    Every answer is an index into the 2,153-module corpus. A module that does not exist cannot come back, so there is no container tag to invent and no include path to get wrong.

  2. 02

    Calibrated confidence

    Each answer carries a probability, so the pipeline can act below a threshold and escalate above it. Confidence is a number the build can branch on.

  3. 03

    2,153 down to 3

    Filtering candidates before the generative step removes about 95% of the prompt tokens an LLM would otherwise read to make the same choice.

  4. 04

    Triage without an API

    Exit 137, exit 143, a missing container, a missing file: diagnosed in 165 ms on the machine that ran the job, offline, at no per-token cost.

Where it runs: Codaris →

05 / Questions

Often asked.

Why not just ask a large model?
A large model can make the same choice, but it has to read the candidate list to do it. Narrowing 2,153 modules to three first removes about 95% of those prompt tokens, answers in 165 ms, and cannot return a module name that does not exist.
Does it generate Nextflow code?
No. It classifies and routes. Code generation stays with a language model, inside Codaris, where every draft is compile-checked and nf-core linted before it lands.
What hardware does it need?
A CPU. It was trained on an Apple Silicon laptop and runs inference there in about 165 ms per question, fully offline.
Can you train one for our domain?
That is the method this model demonstrates: freeze a large encoder, cache its representations, and train a small decision head on a corpus of your decisions. The nf-core run took nine minutes on a laptop once the corpus existed.

Decision models

A decision engine for your domain.

The method behind this model transfers: your corpus, your decisions, a small head on a frozen encoder. Tell us what your team keeps choosing by hand.

Talk through a model →