Vlaander
← All engines

BIO · Life Sciences — Genomics

Bioinformatics Pattern Matching Engine

A dependency-free C++20 exact-match DNA search engine built on an FM-index, whose index files are byte-identical for the same input and embed SHA-256 hashes of the FASTA they were built from.

Maturity. The source you receive (archive 8eb9edd3…02a5, built from Vlaander's repository at commit 19facfb) contains a high-severity memory-safety bug in the Python binding: Index.locate() with a negative max_hits writes past the end of its buffer and crashes or corrupts the host Python process. Vlaander's internal review found it on 7 October 2026, and Vlaander reproduced the crash on the delivered archive on 8 October 2026. The C and C++ APIs are not affected. Two medium findings are also open (see Internal review below).

Overview

BIO counts and locates exact occurrences of DNA patterns in a reference. It uses an FM-index: an SA-IS suffix array, a 2-bit BWT with one 64-byte rank block per 192 bases, and a sampled suffix array for locate. It also aligns reads by exact full-length match on both strands. It does not handle mismatches or gaps.

For the same input, index files are byte-identical whichever compiler or thread count built them. Vlaander showed this with g++ and clang builds and 1 or 4 threads on one x86-64 host. Each index embeds SHA-256 hashes of the input FASTA, each sequence and each section. `bio verify --deep` rebuilds the index from the FASTA and compares it byte for byte. Plain `bio verify` only checks the hashes stored in the index file, which are not signed. It catches accidental corruption, but on its own it does not prove which FASTA an index came from.

A second storage mode, a run-length BWT, grows with the differences between genomes rather than their number. On 16 synthetic 10 Mbp genomes with 0.1% substitutions, its counting structures were 3.8× smaller than the classical mode's. At 5% substitutions they were 1.7× larger. This mode supports counting only.

In Vlaander's review against BWA 0.7.19 and Bowtie 2 v2.5.5 on the same inputs, BIO returned the same exact-hit counts as both. Its exact search ran at about the same speed as `bwa aln`. It built indexes 1.9–3.5× faster, and its indexes were 2.7–3.8× smaller. At 50 Mbp, building took 292 MiB peak memory, against 75 MiB for BWA.

There is no third-party code in the tree. BIO links only the C++ standard library and pthreads. The source is proprietary and contains no copyleft code.

The problem

Regulated genomics pipelines must show which reference an index was built from and that a result can be reproduced. BWA and Bowtie 2 also produced byte-identical indexes on repeated builds in Vlaander's review, but their index files carry no hash of the input FASTA. Both are licensed under GPL-3.0, which limits how they can be shipped inside a closed product. A classical FM-index also grows with the total length of all genomes, even when the genomes are nearly identical.

What it does

  • Exact count and locate over an FM-index (SA-IS suffix array, 2-bit BWT with cache-line rank blocks, sampled suffix array).
  • Two storage modes chosen at build time: classical (count and locate) and a run-length BWT for near-identical genomes (count only).
  • Batched lockstep backward search with prefetching, multi-threaded batch counting, and indexes loaded with mmap, without copying.
  • A documented little-endian index format with SHA-256 hashes of the input FASTA, each sequence and each section. The bytes do not depend on the compiler or thread count.
  • Three interfaces: a C++20 API, a C ABI (libbio.so.1) with opaque handles and status codes, and a Python binding written with ctypes.
  • A command-line tool: build, count, locate, align (exact full-length match on both strands, tab-separated output), verify (with a --deep rebuild check) and info.
  • A bounds-checked index loader, fuzzed with libFuzzer, ASan and UBSan.

Performance

Measured 2026-10-06
MetricFigureBasis
Exact count, 20-mer query, single / batched (three runs)0.64–0.75 / 0.63–0.90 µs at 1 Mbp; 0.73–1.08 / 0.67–0.87 µs at 10 Mbp; 1.34–2.00 / 0.75–1.00 µs at 50 MbpMeasured by Vlaander
Speed-up over counting with std::string::find (single queries, three runs)About 2,800–3,300× at 1 Mbp, 21,000–25,000× at 10 Mbp, 52,000–90,000× at 50 MbpMeasured by Vlaander
Pangenome counting memory vs classical (10 Mbp genomes, 0.1% substitutions)1.24× smaller at 4 genomes, 2.33× at 8, 3.81× at 16; 1.7× larger at 5% substitutionsMeasured by Vlaander
Index reproducibilityByte-identical index files from g++ and clang builds and at 1 or 4 threads, on one x86-64 hostMeasured by Vlaander
Test suite6 ctest suites, 0 failures with g++ and clang; the 5 C++ and CLI suites also pass under ASan+UBSan and TSanMeasured by Vlaander
Date
2026-10-06
Hardware
AMD EPYC 7763, GitHub Codespace (shared Azure VM), 4 vCPU (2 cores), 15 GiB
Toolchain
Rebuild from the delivered archive: g++ 14.3 and clang 16.0.6 with CMake 3.25.1 in a Debian 12 container, Release (-O3), warnings as errors. The builder's own runs (BENCHMARKS.md): g++ 14.2 and clang 14.0.6 in gcc:14-bookworm.
Source archive SHA-256
8eb9edd35e4d7e4c79ce15ecde2fa6c5ae83d312e84523460d94b61ec9c202a5
Method
Vlaander rebuilt the delivered archive and ran the test suites, the sanitizer builds, a g++ versus clang determinism check, bio-bench search and scripts/bench_pangenome.sh. The builder had already run the same benchmarks twice, and the thread-count determinism check once, on the same machine. bio-bench search: uniformly random references of 1, 10 and 50 Mbp in classical mode, 200,000 single-threaded 20-mer queries (half present), and a std::string::find baseline on the first 200 queries, whose counts matched. Pangenome: one random 10 Mbp genome plus copies with the given substitution rate; only the structures needed for counting are compared.
  • Speed figures are ranges over three runs: the rebuild from the archive (the fastest times) and the builder's two runs. The machine is a shared VM with about ±20% run-to-run noise.
  • The std::string::find baseline scans the whole reference for every query, so the speed-up grows with reference size. In Vlaander's review against BWA (7 October 2026), exact search was at parity with bwa aln.
  • The first-edition catalogue's figures (≈2,000× faster, sub-µs independent of reference size, up to 2.5× lower RAM, bit-identical across hosts) are superseded. Single queries can exceed 1 µs at 10 Mbp and do at 50 Mbp. The memory saving needs 8 or more near-identical genomes. Identical output was shown on one host only.
  • Building a pangenome index needs as much memory as building a classical one, because the full suffix array is built first. Only the finished index is smaller.
  • Positions are 32-bit, so a reference can have at most 4,294,967,294 symbols. A human genome of about 3.1 billion bases fits; much larger multi-genome references do not.

Measured figures were produced by Vlaander on the hardware above from the source archive you would receive. The benchmark tools ship with that source, so you can rerun them after delivery; nothing can be run before purchase. Any performance data is illustrative only and not a guarantee of your results. No refund is available after delivery, except where a remedy cannot lawfully be excluded.

Value in your own metrics

Index provenance
Every index embeds SHA-256 hashes of its input FASTA, each sequence and each section. `bio verify --deep` proves an index matches a FASTA by rebuilding it byte for byte. Plain `verify` trusts the unsigned hashes stored in the file.
Memory on pangenome references
Counting structures 2.3–5.0× smaller than classical for 8–32 synthetic near-identical 10 Mbp genomes (0.1% substitutions). About equal at 2% substitutions and 1.7× larger at 5%.
Exact-search speed
About 0.6–2.0 µs per single 20-mer query and 0.6–1.0 µs batched on 1–50 Mbp references, measured by Vlaander on the hardware above on 5–6 October 2026.
Against BWA and Bowtie 2
Vlaander's review, 7 October 2026, single-threaded on the same Codespace pinned to 2 vCPUs: identical exact-hit counts on 400,000 queries; search at parity with bwa aln, about 2× faster than bwa aln plus samse, and about 4.5× faster than Bowtie 2 run as a full aligner; index builds 1.9–3.5× faster and indexes 2.7–3.8× smaller; build peak memory 292 MiB against BWA's 75 MiB at 50 Mbp.
Copyleft exposure
None in the tree. No third-party source is vendored, copied or linked, and the source is proprietary.
Integration surface
A C ABI with opaque handles and status codes, which any language with a C foreign-function interface can call. A Python binding is included.

Assurance and build

Language
C++20, with a C ABI and a Python 3 ctypes binding
Build
CMake 3.20+ · C++ standard library and pthreads only · warnings as errors (-Wall -Wextra -Wpedantic and more); built without warnings with g++ 14 and clang 14 and 16 (checked by Vlaander)
Tests
6 ctest suites: fips (9 SHA-256 tests from FIPS 180-4 and CAVP vectors), unit (18), property (5, with 671,753 assertions comparing count and locate with naive search, batched with single counting, and builds across thread counts), integration (5), a CLI script and Python (5). All pass with g++ and clang (checked by Vlaander, 6 and 8 October 2026)
Sanitizers
The 5 C++ and CLI suites pass under ASan+UBSan and TSan (checked by Vlaander on the delivered archive, 6 October 2026). The Python binding is not run under sanitizers, and the open high finding is in that binding
Fuzzing
libFuzzer + ASan + UBSan with clang 14, builder's runs on 5–6 October 2026: 601 s and 3,081,085 index-loader inputs, 301 s and 725,208 FASTA inputs, no findings. Not re-run by the reviewer
Reproducibility
Byte-identical indexes across g++ and clang builds and across thread counts, on one x86-64 host; not tested on a second host or a big-endian host
Internal review
Vlaander quality gate, 7 October 2026: 0 critical findings. 1 high: the Python locate() overflow, open in the delivered archive. 2 medium, open: plain verify trusts unsigned in-file hashes, and Python close() can unmap the index while another thread is still querying it (reported, not reproduced). 7 low findings reported, not reproduced. There is no CI in the repository, so sanitizer, fuzz and determinism results come from Vlaander's own runs
Source archive SHA-256
8eb9edd35e4d7e4c79ce15ecde2fa6c5ae83d312e84523460d94b61ec9c202a5
Interfaces
C++20 API · C ABI (libbio.so.1) · Python binding (ctypes)
Licensing exposure
Proprietary source; no third-party or copyleft code in the tree (SBOM.md)

Releases are not yet cryptographically signed, and no release manifests or external audit summaries are published. Check the source tarball you receive against its SHA-256 in Schedule 1 of your Sale and Assignment Agreement, which you see before you sign.

Scope and maturity

What is real today, what is deliberately excluded, and what is pending — published unprompted.

Included today

  • The search engine: exact count and locate over the FM-index core, the classical and pangenome storage modes, batched lockstep search, and memory-mapped indexes.
  • Three interfaces (C++20 API, C ABI, Python binding with ctypes) and a CLI (build, count, locate, align, verify, info).
  • The test suites, sanitizer and fuzz harness scripts, the benchmark tool, and README, ARCHITECTURE, BENCHMARKS and SBOM documents.

Explicitly scoped out

  • Not in v1.0: mismatches and indels (alignment is exact full-length match only), paired-end alignment, MAPQ, BAM/CRAM output and gzip input.
  • Not in v1.0: a bidirectional FM-index, long-read, RNA-seq and graph-reference variants, and SIMD or GPU dispatch.
  • References longer than 4,294,967,294 symbols.

Pending

  • The fix for the Python locate() overflow is not yet in the delivered archive.
  • Open medium findings: plain verify does not prove provenance without --deep, and the Python binding has no lock against close() during a query on another thread.
  • Locate in pangenome mode is not implemented, so locate and align need a classical index.
  • Determinism has not been tested on a second host or on a big-endian host. There is no CI.

Who it's for

  • Clinical genomics laboratories that need to show which reference an index was built from.
  • Diagnostics companies embedding exact-match search in a closed commercial product, where GPL-licensed tools are a problem.
  • Research computing groups that count k-mers or exact matches against sets of near-identical genomes.
  • Sequencing instrument makers looking for a small, dependency-free exact-match core.

Before you buy

  1. Check the not-included list against your pipeline first: alignment is exact match only, with no paired-end, MAPQ or BAM output.
  2. Before using the Python binding, fix the locate() overflow yourself or reject a negative max_hits in your own code; the delivered archive does not contain a fix.
  3. For provenance, rely on `bio verify --deep` or an index SHA-256 recorded outside the file, not on plain `verify`.
  4. Reproduce the byte-identical index property on your own machines. Vlaander has shown it on one x86-64 host with two compilers.
  5. Plan build memory: Vlaander's review measured about 6 bytes per base, so a human-scale reference would need very roughly 18–25 GB to build (an estimate, not measured).
  6. Measure the pangenome memory saving on your own references. It needs 8 or more near-identical genomes and disappears as divergence rises.
  7. Check the no-copyleft statement against SBOM.md with your own licence-scanning tooling. Vlaander gives no warranty beyond the limited warranty in the Sale and Assignment Agreement.

What transfers

  • The asset's sourceSubject to the Sale and Assignment Agreement, Vlaander assigns to the verified Buyer the transferable right, title and interest that Vlaander owns in the specified Asset. The assignment excludes Third Party Materials, open-source components, Vlaander’s pre-existing tools, generic know-how, development methods, trademarks, confidential information and any rights that cannot lawfully be assigned.
  • Its testsThe test suite the published assurance claims rest on, as delivered in the source archive.
  • Its audit artefactsSpecifications, model-checking output, reviews and the software bill of materials, where the asset has them.
  • Not includedThird-party and open-source materials are not sold or assigned by Vlaander. They remain under their own licences, listed in each asset’s software bill of materials (SBOM.md in the archive), and the Buyer is responsible for complying with those licences.

First described in: VLA-GEN catalogue, 3 August 2026, §3.6. Revised by Vlaander against the delivered source; every figure is labelled with its basis.

From test to source

  1. 01

    Check

    Read the engine's page: what Vlaander measured and on which hardware, what it has not measured, and the open review findings. Nothing runs before purchase.

  2. 02

    Buy

    Verified businesses only. Place the order, give your company's details, have your authorised signatory sign the Sale and Assignment Agreement, then pay the invoice in naira by bank transfer, or in USDC on Polygon. Once the payment is verified, your account shows Paid.

  3. 03

    Receive

    We confirm the funds and release the source tarball to you. Delivering, then Download ready. Check it against the SHA-256 in Schedule 1, then rerun the tests and benchmarks yourself.