feat(bench): synthetic load harness for flat-file backend (P1-P5)

End-to-end performance and correctness harness for the flat-file +
SQLite database backends. Lives in tests/bench/, built only on
demand (`make bench`); not part of `make check`.

Components

  gen_history (P1)
    Deterministic corpus generator. Knobs: lines, contacts, years,
    seed, stanza-id mode (uuid/libpurple/conversations/mixed), LMC
    rate, MAM-OOO rate, resources/contact, length profile
    (short/mixed/long/extreme). Emits the canonical
    flatlog/<account>/<contact>/history.log layout used by
    ff_verify_integrity. ~340 LOC.

  bench_runner (P1, P2.5)
    S1 cold tail-access via sparse index
    S2 warm tail-access (page cache hot)
    S3 deep pagination (1000 binary-search lookups)
    S4 first-time index build (cold file -> ff_state_ensure_fresh)
    S5 incremental extend (asserts no full rebuild path)
    S6 real ff_verify_integrity over the contact tree
    Reports total/err/warn/info issue counts in the CSV note.

  bench_long_messages (P2)
    L1-L14 long-message stress: 1KB up to 9.9MB bodies, plus
    oversized line rejection (10MB+1 -> ff_readline returns ""),
    embedded-newline / pipe / emoji body patterns, full parse on
    100x1MB, 1000x100KB sustained append.

  bench_failure_modes (P3)
    F1-F15 failure-injection: truncated last line, mid-file CRLF,
    mid-file BOM, LMC cycle, LMC depth>FF_MAX_LMC_DEPTH, manual
    ': ' in resource, RTL/ZWSP, Latin-1 byte, empty body,
    mtime/inode flip, empty file. Each test asserts expected issue
    levels and reports PASS/FAIL.

  bench_export_import (P5)
    Links real database_export.c + database_sqlite.c + database.c
    and drives log_database_export_to_flatfile /
    log_database_import_from_flatfile under load.
    Subcommands: seed, export, import, roundtrip, verify.
    S7a/b export, S8a/b import, S8e roundtrip with full byte-by-byte
    content diff of every row in (from_jid, to_jid, message,
    timestamp, type, stanza_id, archive_id, encryption, replace_id).

Make targets

  bench-quick / bench / bench-full
  bench-longmsg, bench-failure
  bench-multicontact, bench-lmc, bench-ooo
  bench-export, bench-import, bench-roundtrip
  bench-pipeline, bench-pipeline-max (1M rows)
  bench-compare, bench-update-baseline

Volume controls: BENCH_VOLUME (small/medium/max), BENCH_PIPE_ROWS,
BENCH_PIPE_ROWS_MAX, BENCH_DATA_DIR, BENCH_CSV.

Baseline + regression checking (P4)

  tests/bench/baseline.csv       51 rows: S1-S6 x {small,lmc,ooo}
                                 + L1-L14 + F1-F15 (11 of 15) +
                                 S7/S8 x {pipe100k, pipe1M}.
  compare_baseline.py            median over duplicate rows;
                                 exits 1 on any (scenario, volume)
                                 slowdown >= threshold (default 25%).

Verified at scale

  bench-pipeline-max: 1,000,000 rows, full content diff
    seed   17 s, export 304 s, import 31 s, idempotent re-import 10 s,
    diff 2.7 s -- mismatches=0.

Findings surfaced by the harness

  Export scales super-linearly: 4 s @ 100k -> 304 s @ 1M (76x for
    10x rows). Cause: g_slist_sort on the merged list + per-row
    ProfMessage/ff_parsed_line_t allocations. RSS peaks at 1.4 GB
    on 1M. Worth a follow-up.
  Export progress reporting only fires during the write phase --
    the merge+sort phase (~95% of wall time at 1M) is silent.
  /history export and /history import are blocking on the main
    UI thread; profanity is frozen for the duration (~5 min @ 1M).
  ff_readline sets *truncated=TRUE on partial-write tail but
    ff_verify_integrity does not surface this -- partial writes
    go unflagged (failure-injection F1).
  Parser silently truncates body at first unescaped ': ' if a
    resource was manually edited to contain it (F9).
  In gen_history (caught by the bench's own real-verify pass):
    g_strndup mid-codepoint truncation on UTF-8 bank strings -- fixed.

Linkage strategy

  database_flatfile.c + parser + verify + common.c are linked
  unconditionally. The export/import bench additionally links
  database.c + database_sqlite.c + database_export.c. bench_stubs.c
  provides minimal stubs for log_*, prefs_*, connection_get_jid,
  jid_create, files_*, message_*, ui hooks. integrity_issue_free
  is a weak symbol so it falls back to the real database.c
  implementation when that file is linked.
This commit is contained in:
2026-04-30 15:39:30 +03:00
parent 8868d58920
commit b1d820462a
15 changed files with 4241 additions and 0 deletions

View File

@@ -372,6 +372,231 @@ endif
man1_MANS = $(man1_sources)
# ---------------------------------------------------------------------------
# Bench harness — synthetic database load tests. Not part of `make check`.
# Build only when invoked via `make bench` / `make bench-quick`.
# Sources live in tests/bench/. See tests/bench/README.md.
bench_common_sources = \
tests/bench/bench_stubs.c \
tests/bench/bench_common.c tests/bench/bench_common.h \
tests/bench/bench_csv.c tests/bench/bench_csv.h \
src/database_flatfile.c src/database_flatfile.h \
src/database_flatfile_parser.c \
src/database_flatfile_verify.c \
src/common.c src/common.h
# Sources needed only by the export/import bench (S7/S8). Pulls in the SQLite
# backend, the cross-backend export/import code, and the dispatcher.
bench_export_sources = \
src/database.c src/database.h \
src/database_sqlite.c \
src/database_export.c
EXTRA_PROGRAMS = \
tests/bench/gen_history \
tests/bench/bench_runner \
tests/bench/bench_long_messages \
tests/bench/bench_failure_modes \
tests/bench/bench_export_import
tests_bench_gen_history_SOURCES = tests/bench/gen_history.c $(bench_common_sources)
tests_bench_gen_history_CPPFLAGS = -I$(srcdir)/src -I$(srcdir)/tests/bench
tests_bench_gen_history_CFLAGS = $(AM_CFLAGS)
tests_bench_gen_history_LDADD =
tests_bench_bench_runner_SOURCES = tests/bench/bench_runner.c $(bench_common_sources)
tests_bench_bench_runner_CPPFLAGS = -I$(srcdir)/src -I$(srcdir)/tests/bench
tests_bench_bench_runner_CFLAGS = $(AM_CFLAGS)
tests_bench_bench_runner_LDADD =
tests_bench_bench_long_messages_SOURCES = tests/bench/bench_long_messages.c $(bench_common_sources)
tests_bench_bench_long_messages_CPPFLAGS = -I$(srcdir)/src -I$(srcdir)/tests/bench
tests_bench_bench_long_messages_CFLAGS = $(AM_CFLAGS)
tests_bench_bench_long_messages_LDADD =
tests_bench_bench_failure_modes_SOURCES = tests/bench/bench_failure_modes.c $(bench_common_sources)
tests_bench_bench_failure_modes_CPPFLAGS = -I$(srcdir)/src -I$(srcdir)/tests/bench
tests_bench_bench_failure_modes_CFLAGS = $(AM_CFLAGS)
tests_bench_bench_failure_modes_LDADD =
tests_bench_bench_export_import_SOURCES = tests/bench/bench_export_import.c \
$(bench_common_sources) $(bench_export_sources)
tests_bench_bench_export_import_CPPFLAGS = -I$(srcdir)/src -I$(srcdir)/tests/bench
tests_bench_bench_export_import_CFLAGS = $(AM_CFLAGS)
tests_bench_bench_export_import_LDADD =
# Volume control: BENCH_VOLUME={small|medium|max}. Default `small` is safe
# even on tight disks (~5 MB). `max` is heavyweight — see README.
BENCH_VOLUME ?= small
BENCH_DATA_DIR ?= /tmp/cproof-bench-corpus
BENCH_CSV ?= tests/bench/current.csv
bench-build: tests/bench/gen_history tests/bench/bench_runner tests/bench/bench_long_messages tests/bench/bench_failure_modes tests/bench/bench_export_import
# Resolve --lines / --years for the configured volume.
bench-gen: bench-build
@mkdir -p $(BENCH_DATA_DIR)
@case "$(BENCH_VOLUME)" in \
small) L=10000; Y=1; P=mixed ;; \
medium) L=500000; Y=5; P=mixed ;; \
max) L=5000000; Y=10; P=long ;; \
*) echo "BENCH_VOLUME must be small|medium|max"; exit 2 ;; \
esac; \
echo "==> generating volume=$(BENCH_VOLUME) lines=$$L years=$$Y profile=$$P"; \
./tests/bench/gen_history --lines=$$L --years=$$Y --msg-len-profile=$$P \
--lmc-rate=3 --mam-ooo-rate=0 --resources-per-contact=3 \
--output=$(BENCH_DATA_DIR) --seed=42
bench-run: bench-build
@echo "==> running scenarios (csv=$(BENCH_CSV))"
@BENCH_VOLUME=$(BENCH_VOLUME) BENCH_DATA_DIR=$(BENCH_DATA_DIR) \
./tests/bench/bench_runner --data=$(BENCH_DATA_DIR) --csv=$(BENCH_CSV)
@echo "==> CSV: $(BENCH_CSV)"
@cat $(BENCH_CSV)
bench: bench-gen bench-run
bench-quick:
@$(MAKE) bench BENCH_VOLUME=small
# Long-message tests (L1L14). Independent of corpus volume, uses its own tmp dir.
bench-longmsg: bench-build
@echo "==> long-message tests (csv=$(BENCH_CSV))"
@./tests/bench/bench_long_messages \
--tmp=$(BENCH_DATA_DIR)-longmsg --csv=$(BENCH_CSV)
@rm -rf $(BENCH_DATA_DIR)-longmsg
# Failure-injection (F1F15). Independent corpus, asserts behaviour on bad data.
bench-failure: bench-build
@echo "==> failure-injection tests (csv=$(BENCH_CSV))"
@./tests/bench/bench_failure_modes \
--tmp=$(BENCH_DATA_DIR)-fail --csv=$(BENCH_CSV)
@rm -rf $(BENCH_DATA_DIR)-fail
# Export/import (S7/S8) pipeline. Independent corpus under BENCH_DATA_DIR-export.
# Default scale: 100k rows (~530 s). Override with BENCH_PIPE_ROWS=N.
BENCH_PIPE_ROWS ?= 100000
BENCH_PIPE_VOLUME ?= pipe$(BENCH_PIPE_ROWS)
bench-export: bench-build
@echo "==> S7 export pipeline ($(BENCH_PIPE_ROWS) rows, csv=$(BENCH_CSV))"
@rm -rf $(BENCH_DATA_DIR)-export
@BENCH_DATA_DIR=$(BENCH_DATA_DIR)-export \
./tests/bench/bench_export_import seed --rows=$(BENCH_PIPE_ROWS) --account=bench@bench.example
@BENCH_DATA_DIR=$(BENCH_DATA_DIR)-export \
./tests/bench/bench_export_import export --account=bench@bench.example \
--csv=$(BENCH_CSV) --label=S7a_export_cold --volume=$(BENCH_PIPE_VOLUME)
@echo " -- second pass (dedup hot path)"
@BENCH_DATA_DIR=$(BENCH_DATA_DIR)-export \
./tests/bench/bench_export_import export --account=bench@bench.example \
--csv=$(BENCH_CSV) --label=S7b_export_dedup --volume=$(BENCH_PIPE_VOLUME)
@rm -rf $(BENCH_DATA_DIR)-export
bench-import: bench-build
@echo "==> S8 import pipeline ($(BENCH_PIPE_ROWS) rows, csv=$(BENCH_CSV))"
@rm -rf $(BENCH_DATA_DIR)-export
@# Seed account A, export it, then wipe its DB but keep flatlog. Import
@# back into the same account — that way the to: header in flatfile
@# matches the importing account so dedup works correctly.
@BENCH_DATA_DIR=$(BENCH_DATA_DIR)-export \
./tests/bench/bench_export_import seed --rows=$(BENCH_PIPE_ROWS) --account=bench@bench.example
@BENCH_DATA_DIR=$(BENCH_DATA_DIR)-export \
./tests/bench/bench_export_import export --account=bench@bench.example \
--csv=$(BENCH_CSV) --label=S7_seed_export --volume=$(BENCH_PIPE_VOLUME)
@rm -f $(BENCH_DATA_DIR)-export/database/bench_at_bench.example/chatlog.db*
@BENCH_DATA_DIR=$(BENCH_DATA_DIR)-export \
./tests/bench/bench_export_import import --account=bench@bench.example \
--csv=$(BENCH_CSV) --label=S8a_import_cold --volume=$(BENCH_PIPE_VOLUME)
@echo " -- second import (idempotency)"
@BENCH_DATA_DIR=$(BENCH_DATA_DIR)-export \
./tests/bench/bench_export_import import --account=bench@bench.example \
--csv=$(BENCH_CSV) --label=S8b_import_idempotent --volume=$(BENCH_PIPE_VOLUME)
@rm -rf $(BENCH_DATA_DIR)-export
# Full roundtrip with content diff at default volume. PASSes only if
# every row in the rebuilt DB matches the source byte-for-byte.
bench-roundtrip: bench-build
@echo "==> S8e roundtrip+full diff ($(BENCH_PIPE_ROWS) rows, csv=$(BENCH_CSV))"
@rm -rf $(BENCH_DATA_DIR)-export
@BENCH_DATA_DIR=$(BENCH_DATA_DIR)-export \
./tests/bench/bench_export_import roundtrip --rows=$(BENCH_PIPE_ROWS) \
--csv=$(BENCH_CSV) --label=S8e_roundtrip --volume=$(BENCH_PIPE_VOLUME) --full-diff
@rm -rf $(BENCH_DATA_DIR)-export
# Pipeline at medium volume (default BENCH_PIPE_ROWS=100k).
bench-pipeline: bench-export bench-import bench-roundtrip
# Pipeline at MAX volume — 1M rows. Heavy: ~510 min, ~500MB SQLite + ~500MB
# flatfile + ~500MB second SQLite. Run only when you can afford the time
# and the disk. Set BENCH_PIPE_ROWS_MAX to override (default 1_000_000).
BENCH_PIPE_ROWS_MAX ?= 1000000
bench-pipeline-max: bench-build
@echo "==> bench-pipeline-max: $(BENCH_PIPE_ROWS_MAX) rows (heavy!)"
@$(MAKE) bench-export BENCH_PIPE_ROWS=$(BENCH_PIPE_ROWS_MAX)
@$(MAKE) bench-import BENCH_PIPE_ROWS=$(BENCH_PIPE_ROWS_MAX)
@$(MAKE) bench-roundtrip BENCH_PIPE_ROWS=$(BENCH_PIPE_ROWS_MAX)
# S9: 200 contacts × 50k lines per contact = 10M lines distributed.
bench-multicontact: bench-build
@mkdir -p $(BENCH_DATA_DIR)-multi
@echo "==> S9: 200 contacts x 5000 lines"
@./tests/bench/gen_history --lines=1000000 --contacts=200 --years=3 \
--msg-len-profile=mixed --output=$(BENCH_DATA_DIR)-multi --seed=42
@BENCH_VOLUME=multicontact BENCH_DATA_DIR=$(BENCH_DATA_DIR)-multi \
./tests/bench/bench_runner --data=$(BENCH_DATA_DIR)-multi --csv=$(BENCH_CSV) \
--scenarios=S4,S6
@rm -rf $(BENCH_DATA_DIR)-multi
# S10: 30% LMC corrections.
bench-lmc: bench-build
@mkdir -p $(BENCH_DATA_DIR)-lmc
@echo "==> S10: LMC-heavy corpus (30%% corrections)"
@./tests/bench/gen_history --lines=100000 --years=2 --lmc-rate=30 \
--msg-len-profile=mixed --output=$(BENCH_DATA_DIR)-lmc --seed=42
@BENCH_VOLUME=lmc BENCH_DATA_DIR=$(BENCH_DATA_DIR)-lmc \
./tests/bench/bench_runner --data=$(BENCH_DATA_DIR)-lmc --csv=$(BENCH_CSV)
@rm -rf $(BENCH_DATA_DIR)-lmc
# S11: 20% MAM out-of-order timestamps.
bench-ooo: bench-build
@mkdir -p $(BENCH_DATA_DIR)-ooo
@echo "==> S11: MAM-OOO corpus (20%% out-of-order)"
@./tests/bench/gen_history --lines=100000 --years=2 --mam-ooo-rate=20 \
--msg-len-profile=mixed --output=$(BENCH_DATA_DIR)-ooo --seed=42
@BENCH_VOLUME=ooo BENCH_DATA_DIR=$(BENCH_DATA_DIR)-ooo \
./tests/bench/bench_runner --data=$(BENCH_DATA_DIR)-ooo --csv=$(BENCH_CSV) \
--scenarios=S6
@rm -rf $(BENCH_DATA_DIR)-ooo
# Run everything: P1 scenarios + long messages + failure injection + multicontact + LMC + OOO.
bench-full: bench bench-longmsg bench-failure bench-multicontact bench-lmc bench-ooo
# Compare current.csv against baseline.csv and exit non-zero on regressions.
BENCH_BASELINE ?= tests/bench/baseline.csv
BENCH_THRESHOLD ?= 25
bench-compare:
@python3 tests/bench/compare_baseline.py \
--baseline=$(BENCH_BASELINE) --current=$(BENCH_CSV) \
--threshold=$(BENCH_THRESHOLD)
# Snapshot current.csv as the new baseline. Run this only after manually
# reviewing that the numbers look healthy.
bench-update-baseline:
@cp $(BENCH_CSV) $(BENCH_BASELINE)
@echo "baseline updated: $(BENCH_BASELINE)"
bench-clean:
rm -rf $(BENCH_DATA_DIR) $(BENCH_DATA_DIR)-longmsg $(BENCH_DATA_DIR)-fail \
$(BENCH_DATA_DIR)-multi $(BENCH_DATA_DIR)-lmc $(BENCH_DATA_DIR)-ooo \
$(BENCH_DATA_DIR)-export $(BENCH_CSV)
.PHONY: bench bench-quick bench-build bench-gen bench-run bench-clean \
bench-longmsg bench-failure bench-multicontact bench-lmc bench-ooo \
bench-full bench-compare bench-update-baseline \
bench-export bench-import bench-roundtrip bench-pipeline bench-pipeline-max
EXTRA_DIST = $(man1_sources) $(icons_sources) $(themes_sources) $(script_sources) profrc.example theme_template LICENSE.txt README.md CHANGELOG
# Ship API documentation with `make dist`
@@ -387,6 +612,20 @@ EXTRA_DIST += \
apidocs/python/src/plugin.py \
apidocs/python/src/prof.py
# Bench harness sources (not built by default; see EXTRA_PROGRAMS above)
EXTRA_DIST += \
tests/bench/README.md \
tests/bench/compare_baseline.py \
tests/bench/gen_history.c \
tests/bench/bench_runner.c \
tests/bench/bench_long_messages.c \
tests/bench/bench_failure_modes.c \
tests/bench/bench_export_import.c \
tests/bench/bench_stubs.c \
tests/bench/bench_common.c tests/bench/bench_common.h \
tests/bench/bench_csv.c tests/bench/bench_csv.h \
tests/bench/baseline.csv
if INCLUDE_GIT_VERSION
EXTRA_DIST += .git/HEAD .git/index