Performance
Every number on this page comes from docs/benchmark-results/ in the celeriant-db repo: ec2-benchmark.md and its CSV (EC2, Sydney ap-southeast-2, August and September 2026, main at the time), rpi-benchmark.md, kafka-benchmark.md and marten-benchmark.md. Each figure carries the setup it was measured under. A number without its setup is useless, so none appear here.
How throughput works
The load generator runs N tasks. Each task holds one connection, writes one event to its own aggregate, waits for the ack, then writes the next. One request in flight per connection, no batching, no pipelining.
So throughput is connections divided by latency. Nothing else. More connections buy more throughput only while latency stays flat; past the knee, latency grows faster than connections and throughput falls.
Latency has a floor, and under load that floor is the amortisation windows. A write waits for its batch's fsync window, then its fsync, and in a cluster the replication window and the round trip to the follower, on both nodes, before the ack. The quiet-server run on the flagship shows it plainly:
| connections | writes/s | p50 | p99 | fsync window | replication window |
|---|---|---|---|---|---|
| 8 | 433 | 18 ms | 19 ms | 1,000 µs | 15,000 µs |
| 64 | 3,446 | 18 ms | 19 ms | 1,000 µs | 15,000 µs |
2x i4i.metal, replicated, mTLS, 2026-09-06. The CSV note attributes the 18 ms to the fsync window plus the replication window plus the network round trip, and records that the coordinator's idle fast path (which should skip the fsync window on an idle shard) was not engaging in that run. At 64 connections the box is under 1% busy. It is waiting, not working.
The consequence: a small client pool on a big server measures the windows, not the server. To find the ceiling you need tens of thousands of concurrent writers.
Pick a tier
Peak clean throughput: the highest level with zero client-visible errors. All replicated tiers run mTLS on both client and replication paths.
| tier | cluster | storage | writes/s | connections | p50 | p99 | fsync / repl window | stat | cost |
|---|---|---|---|---|---|---|---|---|---|
| Flagship | 2x i4i.metal (128 vCPU) | 8x NVMe RAID0 | 1,057,417 | 60,000 | 48 ms | 108 ms | 1,000 / 15,000 µs | worst of 3 | ~$3,700/mo spot |
| Large | 2x i4i.16xlarge (64 vCPU) | 4x NVMe RAID0 | 466,640 | 60,000 | 97 ms | 264 ms | 17,000 / 17,000 µs | single run | ~$9,600/mo |
| Mid | 2x i4i.8xlarge (32 vCPU) | 2x NVMe RAID0 | 446,667 | 24,000 | 44 ms | 109 ms | 4,000 / 8,000 µs | worst of 3 | ~$4,800/mo |
| Entry | 2x c7g.xlarge (4 vCPU ARM) | gp3, 3,000 IOPS | 67,720 | 8,000 | 113 ms | 165 ms | 17,000 / 17,000 µs | worst of 3 | ~$295/mo |
| Single node | 1x c6i.xlarge (4 vCPU) | gp3, 3,000 IOPS | 86,658 | 8,000 | 85 ms | 108 ms | 17,000 µs / n.a. | worst of 3 | ~$172/mo |
| Floor | 2x c6g.medium (1 vCPU) | gp3, 3,000 IOPS | 12,289 | 3,000 | 183 ms | 284 ms | not recorded | single run | ~$82/mo |
Cost is Sydney on-demand at 730 h/month plus gp3 volumes, except the metal pair, quoted at the $5.08/hr spot price the sweep ran on. On demand it is $26.33/hr, about $19,200 a month. The mid-tier windows were tuned on that instance; the 17,000 µs rows ran on the fsync default of the time, which has since changed (see Tuning).
The 32-vCPU pair gets 96% of the 64-vCPU pair's throughput for half the money. Cores that share an SMT sibling are not cores; 64 physical cores on the metal box is where the next step comes from.
The single node beats the replicated ARM pair because it skips the replication round trip. On EBS the volume outlives the instance, so that is a defensible trade: the second node buys availability, not durability.
Per dollar, the entry pair wins: 230 writes/s per dollar per month, against 55 for the metal pair at on-demand prices.
The flagship ladder
2x i4i.metal (64 physical cores each), 8x NVMe RAID0, 128 shards, fsync window 1,000 µs, replication window 15,000 µs, four c6i.4xlarge load generators, single AZ, 2026-08-13. Worst of three repetitions.
| connections | writes/s | p50 | p95 | p99 | CPU busy | run-to-run spread |
|---|---|---|---|---|---|---|
| 32,000 | 813,019 | 36 ms | 43 ms | 60 ms | 82% | 1.2% |
| 60,000 | 1,057,417 | 48 ms | 72 ms | 108 ms | 87% | 0.8% |
| 100,000 | 1,018,950 | 77 ms | 146 ms | 220 ms | 94% | 5.3% |
| 132,000 | 916,122 | 112 ms | 206 ms | 432 ms | 95% | 10.1% |
60,000 is the peak and the most reproducible point. Past it, CPU pins, throughput falls, and the spread widens. The 100,000 and 132,000 rows also logged errors (571 and 575).
The same box standalone and in cleartext does 1,936,064 writes/s at 32,000 connections. Replication plus mTLS costs about 45% of that.
Storage
Same instance (2x i4i.16xlarge), same clients; only the storage differs. gp3 is provisioned to its maximum, 16,000 IOPS and 1,000 MB/s. 2026-08-10.
| connections | NVMe RAID0 | p99 | gp3 | p99 |
|---|---|---|---|---|
| 24,000 | 249,393 | 243 ms | 253,362 | 480 ms |
| 48,000 | 401,838 | 250 ms | 313,506 | 772 ms |
| 72,000 | 407,225 | 251 ms | 233,547 | 1,124 ms |
| 132,000 | 403,201 | 258 ms | 148,542 | 1,211 ms |
gp3 keeps up at low load, peaks, then declines. NVMe holds flat. The EBS path is latency-bound, not IOPS-bound: the volume drew about 2,100 to 2,600 write IOPS of the 16,000 provisioned. Buying more IOPS buys nothing. At 120,000 connections on gp3 the leader could not renew its lease, shards were fenced and 667 writes failed before the cluster recovered. NVMe took zero errors at every level.
Stripe every drive. At 64 vCPU, RAID0 across four drives was +32% throughput over one.
Tuning
Two windows amortise work across concurrent writes: the fsync window (--fsync-delay-us, default 4000) and the replication window (--replication-delay-us, default 17000). See Configuration.
- Fsync window: keep it under about 6 ms. On i4i.metal standalone at 32,000 connections, everything from 100 µs to 6,400 µs landed within run-to-run noise. The old 17,000 µs default cost 33.2%. On gp3, 17,000 µs cost 29.6%. The lost time shows as iowait, and p99 flatters the bad setting because less throughput means less queueing. Compare medians.
- Replication window: do not copy between instance types. 15,000 µs was the metal optimum. On i4i.8xlarge the same value was 34.8% worse. It lives in the
i4i-metaldeploy profile, not in the defaults. - Shard count defaults to the vCPU count, which wins near saturation (+13.3% at 32,000 connections on metal) and loses badly at moderate load (three times slower at 8,000 connections).
- Small boxes: the fsync window on a 4-vCPU node was worth at most +2.4%. CPU binds first. Ship them stock.
Where it breaks
Past its knee a tier does not degrade gracefully.
- The c7g.xlarge pair runs clean to 8,000 connections and sheds 70,694 errors at 16,000 (worst of three), p99 2.5 s.
- The c6g.medium pair collapsed above 4,000 connections, went to zero throughput at 8,000, and stayed wedged at 99.5% CPU for 15+ minutes after the clients left. Open robustness issue.
Know your tier's ceiling and stay under it.
Against Kafka and PostgreSQL
Both ran on i4i.8xlarge in March 2026 with the same harness shape: N tasks, one write each, wait for the ack. Neither ran at the settings above, so read the setups.
- Kafka 4.0.2 (KRaft): 3 brokers, 2x c7i.4xlarge clients,
acks=all,replication.factor=2,min.insync.replicas=2, batching off, TLS. No fsync: acknowledged from page cache. - Marten 7 on PostgreSQL 17.8: primary plus synchronous standby, 3x c7i.4xlarge clients,
synchronous_commit = on, mTLS. OneEvents.Append+SaveChangesAsyncper event. Both nodes fsync WAL before ack.
| connections | Kafka | Marten | Celeriant (comparison column) |
|---|---|---|---|
| 500 | 42,721 (peak) | ||
| 9,000 | 17,491 | 12,666 | 144,655 |
| 12,000 | 19,559 | 901 | 190,647 |
| 24,000 | 21,212 | 1,651 | 318,768 |
| 60,000 | 24,162 (peak) | 17,513, 36,624 errors |
The Celeriant column is the one stored beside Kafka's and Marten's in their CSVs, on 2x i4i.8xlarge. The files do not record its fsync or replication windows. It errored from 48,000 connections up in that run. The Marten CSV's Celeriant values below 9,000 connections are round numbers (80,000, 105,000 and so on) and are left out here as not measured.
Marten peaks at 500 connections, then hits the process-per-connection wall: from 12,666 writes/s at 9,000 connections to 901 at 12,000. Kafka is flat and never fast. Its p99 is 1,342 ms at 24,000 connections.
Current Celeriant on the same instance type is the mid tier above: 446,667 writes/s at 24,000 connections, p99 109 ms, measured in August, not side by side.
Raspberry Pi 5
2x Raspberry Pi 5 (4 GB, 4x Cortex-A76), one NVMe M.2 SSD each, XFS, gigabit LAN, one client machine, 60 s per level, replicated over mTLS with kTLS, 2026-03-25. Fsync window not recorded.
Peak 35,382 writes/s at 10,000 connections, avg 280.7 ms, p99 462 ms, zero errors from 250 to 10,000 connections. Errors start at 12,000 (client-side timeouts). The sweet spot is 250 to 2,000 connections: avg 67 to 93 ms, p99 under 110 ms.
Method
- Workload:
rpi_cluster_pool_bench, one event per acknowledged write. Every write isfdatasynced through Direct I/O on both nodes (XFS, io_uring) and acknowledged after both succeed. Replication runs over mTLS with kTLS offload. The benchmark files do not record the payload size; the harness writes a short text payload of the form[t-<task>-r-<seq>] hello. - Worst of N: throughput is the minimum across repetitions, latency the maximum, cells run forward then backward so drift cancels. The CSV's
statcolumn says where only one repetition exists. - One connection pinned per writer task. A shared pool lets connections drift between aggregates on different shards and inflated p99 up to 40x in testing. That is the harness, not the database.
- Rejected runs: zero throughput, replication not established, load shedding, missing load generators.
What it does not cover: reads, mixed load, large events, and failure paths. Every cell is a clean cluster in a single AZ. Expect cross-AZ to be slower. The reproduction commands are in ec2-benchmark.md. See Durability and safety for what an ack guarantees.