Skip to content

nf-pilot · Released · MIT

The fast first answer for your pipeline agent.

nf-pilot answers the routine Nextflow setup questions on your own machine in about 0.6 seconds, with no API call and no tokens. It chooses from what already exists in nf-core, so it cannot invent a module, and every answer comes with a confidence score.

0
LLM tokens per decision

runs locally, no API

~0.6 s
Per decision on a laptop CPU

about 45 ms on an Apple Silicon GPU

2,153
nf-core modules it can choose from

a fixed list, nothing invented

01 / What it decides

The questions it answers.

Setting up a pipeline is a run of small choices. nf-pilot makes them the way an experienced nf-core developer would, and says how sure it is.

Which module does this step need?
It picks from 2,153 nf-core modules and their containers, so the answer is always one that exists.
Nanopore or PacBio reads?
It swaps short-read QC for a long-read tool: NanoPlot in place of FastQC, for example.
Is there already a pipeline for this assay?
It matches the request to a released nf-core pipeline before anyone builds a new one.
How much memory and time?
It sets the resources and retry rules a process should start with.
One step, or a subworkflow?
It knows when a chain of steps should be packaged together, and when an official nf-core subworkflow already does the job.
Is this good Nextflow?
It tells idiomatic DSL2 from the common anti-patterns before they reach a run.

02 / Against an LLM

Fast and free on its own. Most accurate alongside an LLM.

We gave 40 pipeline setup decisions to nf-pilot, to a frontier LLM, and to the two working together.

nf-pilotLLM aloneBoth together
Correct, out of 40283234
Tokens used05,4337,872
Time per decision0.6 s2.0 s2.8 s

Ten tasks each in container choice, subworkflow packaging, read-type QC and samplesheet checks. Figures from the nf-pilot model card.

Where it is still weak: samplesheet checks. nf-pilot gets 5 of 10 right there and the LLM gets 4, so have a person review those.

03 / Updates

It gets better with each update.

The latest update taught nf-pilot four new kinds of decision and nearly doubled what it learned from.

At launch Now
  • Decisions answered correctly69.4 → 78.4%
Figure 1 | Share of 1,165 check questions answered correctly, at launch and now. Same questions both times.
  1. Latest updateFour new decision types: read-type QC, resource settings, subworkflow packaging and samplesheet checks. Training examples up from 6,082 to 11,654.
  2. First releaseFirst release: module choice, pipeline matching, DSL2 checks and request routing.

04 / How it works

It chooses. It never writes.

nf-pilot does not generate text. It reads the request and picks one option from a fixed list. That is why it runs on a laptop, and why it cannot make up a module name. Writing the code stays with your LLM.

your request→nf-pilot picks from a fixed list→answer + confidence

Technical details on the model card ↗

05 / Questions

Often asked.

Does it replace my LLM?
No. It takes the quick, repeatable choices. In our 40-task comparison the LLM alone was still more accurate, 32 against 28, and the two together did best at 34.
Does it write Nextflow code?
No. It chooses and classifies. Code generation stays with a language model; inside Codaris, every draft is checked before it lands.
What does it need to run?
A laptop. About 0.6 seconds per decision on a CPU, about 45 ms on an Apple Silicon GPU, fully offline.
Can you build one for our domain?
Yes. The approach works on any set of repeatable decisions: we collect examples of the choices your team makes and train a small model on them. Training nf-pilot took under half an hour on a laptop once the examples existed.

Custom models

A decision model for your domain.

Tell us what your team keeps choosing by hand. We will tell you whether a small model can take it over.

Talk through a model →