Skip to content

§ Practice Directory · 5 Core Competencies

Five practices, one accountable principal.

Each practice stands alone. Together they take a project from the sequencer to a model running on a laptop, with the same person accountable end to end.

§ 01Practice

Omics Analysis

Analysis you can defend in a review.

WGS / WESBulk & single-cell RNA-seqATAC / ChIP / methylationMetagenomicsMulti-omics

Practice Overview

We take raw sequencing, mass-spec or array data through QC, processing, statistics and interpretation. Every result comes with the methods, the code and the environment needed to reproduce it.

Core Capabilities

  • Genomics

    Germline and somatic variant calling, structural variants, CNVs, annotation and prioritisation. Short-read and long-read (ONT, PacBio).

  • Transcriptomics

    Bulk RNA-seq differential expression, splicing and fusion detection. Single-cell and spatial: clustering, annotation, trajectories, cell–cell communication.

  • Epigenomics

    ATAC-seq, ChIP-seq and CUT&Tag peak calling, differential accessibility, motif analysis. Bisulfite and nanopore methylation.

  • Proteomics & metabolomics

    Label-free and TMT quantification, imputation, differential abundance, pathway enrichment.

  • Microbiome & metagenomics

    Taxonomic profiling, assembly and binning, functional annotation, AMR gene detection, diversity statistics.

  • Multi-omics integration

    Factor models, network integration and joint embeddings that connect layers instead of stapling results together.

Key Deliverables

  • Analysis-ready matrices, variant tables and annotated objects
  • A written report with methods you can paste into a manuscript
  • Publication-quality static figures and interactive views
  • Reproducible notebooks with pinned environments
§ 02Practice

Pipeline Engineering

Pipelines that still run in five years.

Nextflow / nf-coreSnakemakeWDLContainersHPC & cloud

Practice Overview

A pipeline is software. We build yours with version control, tests, containers and CI, so it runs the same way for every sample, on every machine, and can be handed to the next person without a walkthrough.

Core Capabilities

  • New pipelines

    Nextflow to nf-core standards, Snakemake or WDL, depending on your team and infrastructure. Modular, documented, parameterised.

  • Legacy migration

    Turn a folder of Bash, R and Python scripts into a tested workflow without changing the science it encodes.

  • Containers & environments

    Docker and Apptainer images, Conda and pixi lockfiles, and a strategy for keeping them current.

  • Execution anywhere

    SLURM, PBS, LSF on-prem; AWS Batch, Google Batch and Azure in the cloud; Seqera Platform for orchestration.

  • Testing & CI

    nf-test and pytest suites, small test datasets, GitHub Actions or GitLab CI that run on every change.

  • Cost & performance

    Runtime and cost profiling, resource right-sizing, spot and retry strategies that cut cloud bills without losing runs.

Key Deliverables

  • A versioned pipeline repository with CI, tests and docs
  • Container images and resolved, reproducible environments
  • Runbooks for your cluster or cloud account
  • Runtime and cost benchmarks on your data
§ 03Practice

Agentic RAG Systems

Answers with sources, not vibes.

Hybrid retrievalLive database toolsCitationsEvaluationOn-prem

Practice Overview

Retrieval-augmented generation done properly: your documents, databases and live APIs behind an agent that plans, retrieves, checks its own work and cites what it used. Runs on your infrastructure or entirely offline.

Core Capabilities

  • Ingestion

    PDFs, ELN and LIMS exports, internal reports, sequence and variant databases, wikis and tickets. Chunking and metadata that match how scientists search.

  • Hybrid retrieval

    Dense embeddings plus BM25 with reranking, tuned on your own query logs. Domain embeddings for gene, variant and compound names.

  • Agentic loops

    The agent decides when to query PubMed, Ensembl, UniProt, ClinVar or your SQL, and when to stop. Multi-step questions get multi-step plans.

  • Grounded answers

    Every claim links to a source span. Unsupported claims are flagged, not smoothed over.

  • Evaluation

    Retrieval recall, faithfulness and answer quality measured on a test set built with your team, before and after every change.

  • Deployment

    Containerised service with an API, access control and audit logs. Cloud, on-prem or fully offline with an on-device model.

Key Deliverables

  • A retrieval service and API you own
  • An evaluation suite with baseline numbers
  • Ingestion pipelines that keep the index fresh
  • An admin interface for sources and permissions
§ 04Practice

AI Agents

Agents that do the work, and show their work.

Pipeline operatorsQC triageReport draftingMCP toolsGuardrails

Practice Overview

We build agents that operate inside your existing tools: they launch pipelines, read QC reports, curate literature, draft documents and hand off to a person at the points you choose. Every action is logged and reversible.

Core Capabilities

  • Analysis operators

    Agents that launch Nextflow runs, watch for failures, retry sensibly and summarise results when a run completes.

  • QC triage

    Read MultiQC and pipeline metrics, flag outlier samples, propose re-runs, and explain why in plain language.

  • Curation assistants

    Literature screening, variant curation drafts, and structured extraction from papers into your schema.

  • Report drafting

    Structured outputs into your templates: methods, results, figure captions, with links back to the data.

  • Tool integration

    Model Context Protocol servers for your LIMS, ticketing, Slack, cluster and cloud, with scoped permissions.

  • Safety & oversight

    Approval gates, budgets, rate limits, full trace logs, and evaluation and red-teaming before anything touches production.

Key Deliverables

  • Agents with defined tools, permissions and escalation paths
  • Trace logs and dashboards for every action taken
  • An evaluation suite and a staged rollout plan
  • Training for the team that will run them
§ 05Practice

On-device AI Models

Frontier-quality answers, laptop-sized models.

~3B parametersFine-tuningMLX / llama.cpp / Core MLOfflinemacOS apps

Practice Overview

For narrow, well-defined tasks, a small model trained on the right data beats a large general one, and it runs on the MacBook already on your desk. We select, fine-tune, quantise and package models that work fully offline.

Core Capabilities

  • Model selection

    Baseline evaluations across open model families (Qwen, Llama, Gemma, Phi and others) on your actual task, with licensing checked.

  • Fine-tuning

    Supervised and preference tuning with LoRA and full fine-tunes; distillation from frontier models when you lack labelled data.

  • Quantisation & packaging

    4-bit and 8-bit quantisation for MLX, llama.cpp and Core ML. Accuracy loss measured, not assumed.

  • Native deployment

    Signed macOS apps, command-line tools or a local API server. Works on a plane, in a clean room, or behind an air gap.

  • Profiling

    Tokens per second, memory footprint, thermal and battery behaviour on the M-series chips your team actually uses.

  • Lifecycle

    An evaluation and update pipeline so the model improves as your data does, without regressions.

Key Deliverables

  • Model weights, the training recipe and an evaluation report
  • A signed macOS app or CLI, or a local server
  • Fine-tuning and evaluation code you own
  • An update playbook for the next version

§ Practice Scoping

Scope a dataset, pipeline audit, or custom model.

Send a description of your assay types, sample numbers, and deadlines. We reply with a preliminary scoping memo within two working days.

Scope Your Project →