Performance¶
How fast MyVector's HNSW search is, and how much accuracy you trade for speed.
Measured on MyVector (plugin, MySQL 8.4) with 100,000 GloVe 6B 50d vectors and 1,000 queries that are not in the index (Cosine distance, k = 10), on 2026-09-29 at commit 4b0789f. Host: 8-core Arm Neoverse-N1 (aarch64), 46 GB RAM.
All numbers on this page come from
docs/data/performance.json.
Recall vs throughput¶
Each point is one ef_search setting: the lowest is the fastest, the highest the most
accurate. Recall@10 is the share of the true 10 nearest neighbours that the search returns.
| ef_search | recall@10 | QPS | p50 ms | p99 ms |
|---|---|---|---|---|
| 10 | 0.809 | 1,204 | 0.8 | 1.0 |
| 20 | 0.907 | 1,124 | 0.9 | 1.1 |
| 50 | 0.974 | 1,002 | 1.0 | 1.2 |
| 100 | 0.992 | 832 | 1.2 | 1.4 |
| 200 | 0.999 | 612 | 1.7 | 1.9 |
| 400 | 1.000 | 481 | 2.1 | 2.6 |
ef_search is how many candidate neighbours the HNSW search keeps while it walks the index
graph. Higher values find more of the true nearest neighbours (higher recall) but take longer
per query. Set it per query with MYVECTOR_IS_ANN(index, key, vector, 'nn=10,ef_search=N');
queries without it use the index's own setting. QPS here is measured through the mysql
client in Docker, one query at a time, so treat it as relative between ef_search values
(#133).
Latest release¶
Release v1.26.9 in CI (synthetic, 10,000 rows × 128 dimensions, GitHub-hosted runners): index build and insert throughput for each supported MySQL version and build.
| MySQL / build | Index build | Insert QPS | recall@10 |
|---|---|---|---|
| 8.4 plugin | 2.95 s | 3,707 | 0.978 |
| 8.4 component | 2.67 s | 3,942 | — |
| 9.7 component | 2.64 s | 3,916 | — |
| 26.7 component | 2.53 s | 4,062 | — |
"—" means no recall was measured: component builds from v1.26.9 and earlier don't have the
query rewrite that MYVECTOR_IS_ANN needs (#156).
Search QPS from these CI runs isn't shown, because up to v1.26.9 it mostly measured Docker
overhead (#124).
Comparison with MariaDB (2025)
A one-off comparison with MariaDB on the ann-benchmarks gist-960-euclidean and
dbpedia-openai-1000k-angular datasets, run in February 2025 on a pre-1.0 build. It used
different data, hardware and settings, so its numbers can't be compared with the ones
above. See MariaDB Comparison (2025).
Test environment and method
- Queries: sent one at a time. They are held out from the dataset by a seeded shuffle and never inserted, so no query finds itself.
- Index: HNSW, M = 16, ef_construction = 200. Recall is measured against an exact brute-force top 10.
- Release results: the
myvectorbenchworkflow on eachv*tag, stored on thebenchmarksbranch.
To reproduce the sweep:
python3 scripts/myvectorbench.py --config myvectorbench-glove.yml \
--mysql-version 8.4 --build-path plugin --artifact-dir dist/plugin-8.4 \
--image ghcr.io/askdba/myvector:mysql8.4 --output glove-sweep.json
Full write-up, including both runs and their latency: ef_search Sweep.