nf-pilot · Released · MIT
The fast first answer for your pipeline agent.
nf-pilot answers the routine Nextflow setup questions on your own machine in about 0.6 seconds, with no API call and no tokens. It chooses from what already exists in nf-core, so it cannot invent a module, and every answer comes with a confidence score.
- 0
- LLM tokens per decision
- ~0.6 s
- Per decision on a laptop CPU
- 2,153
- nf-core modules it can choose from
runs locally, no API
about 45 ms on an Apple Silicon GPU
a fixed list, nothing invented
01 / What it decides
The questions it answers.
Setting up a pipeline is a run of small choices. nf-pilot makes them the way an experienced nf-core developer would, and says how sure it is.
- Which module does this step need?
- It picks from 2,153 nf-core modules and their containers, so the answer is always one that exists.
- Nanopore or PacBio reads?
- It swaps short-read QC for a long-read tool: NanoPlot in place of FastQC, for example.
- Is there already a pipeline for this assay?
- It matches the request to a released nf-core pipeline before anyone builds a new one.
- How much memory and time?
- It sets the resources and retry rules a process should start with.
- One step, or a subworkflow?
- It knows when a chain of steps should be packaged together, and when an official nf-core subworkflow already does the job.
- Is this good Nextflow?
- It tells idiomatic DSL2 from the common anti-patterns before they reach a run.
02 / Against an LLM
Fast and free on its own. Most accurate alongside an LLM.
We gave 40 pipeline setup decisions to nf-pilot, to a frontier LLM, and to the two working together.
| nf-pilot | LLM alone | Both together | |
|---|---|---|---|
| Correct, out of 40 | 28 | 32 | 34 |
| Tokens used | 0 | 5,433 | 7,872 |
| Time per decision | 0.6 s | 2.0 s | 2.8 s |
Ten tasks each in container choice, subworkflow packaging, read-type QC and samplesheet checks. Figures from the nf-pilot model card.
Where it is still weak: samplesheet checks. nf-pilot gets 5 of 10 right there and the LLM gets 4, so have a person review those.
03 / Updates
It gets better with each update.
The latest update taught nf-pilot four new kinds of decision and nearly doubled what it learned from.
- Decisions answered correctly69.4 → 78.4%
- Latest updateFour new decision types: read-type QC, resource settings, subworkflow packaging and samplesheet checks. Training examples up from 6,082 to 11,654.
- First releaseFirst release: module choice, pipeline matching, DSL2 checks and request routing.
04 / How it works
It chooses. It never writes.
nf-pilot does not generate text. It reads the request and picks one option from a fixed list. That is why it runs on a laptop, and why it cannot make up a module name. Writing the code stays with your LLM.
05 / Questions
Often asked.
- Does it replace my LLM?
- No. It takes the quick, repeatable choices. In our 40-task comparison the LLM alone was still more accurate, 32 against 28, and the two together did best at 34.
- Does it write Nextflow code?
- No. It chooses and classifies. Code generation stays with a language model; inside Codaris, every draft is checked before it lands.
- What does it need to run?
- A laptop. About 0.6 seconds per decision on a CPU, about 45 ms on an Apple Silicon GPU, fully offline.
- Can you build one for our domain?
- Yes. The approach works on any set of repeatable decisions: we collect examples of the choices your team makes and train a small model on them. Training nf-pilot took under half an hour on a laptop once the examples existed.
Custom models
A decision model for your domain.
Tell us what your team keeps choosing by hand. We will tell you whether a small model can take it over.
Talk through a model →