PQC Verified

Eight months of building a C web framework

by gg582 · 2026-09-22 11:28:42 · 101 views · 10 min read

Table of contents

Hi, We are Team C 4 Punk Developers. We are introducing our project called CWIST.

CWIST is a web framework written in C. It serves HTTP/1.1, HTTP/2, and HTTP/3. It has WebSockets, SSE, GraphQL, a gRPC client with real retry logic, SQLite with migrations and pooling, and it compiles to WebAssembly, both for the browser and for WASI. There are 86 test files, 31 tutorials, and a Homebrew tap.

The first commit was on January 7, 2026. About 1,200 commits have landed since. This post is a look back at that period. The failures stay in, because the failures are the interesting part.

The start

The project was originally called CWISTAW, short for "C Web development Is STill Alive and Well." Two days later it was renamed to CWIST: "C Web development Is Still Trustworthy." The rename changed two lines in the README. What survived from those first days: a 140-line http.c, a small string library, an error-code module, and cJSON as the only dependency.

Two things were already true on day three. The server had been run under ab, and it already had its first stress-test crash fix. Also, the original http.c forced a multiprocess model, and that was ripped out within 48 hours. That tension, getting fast numbers under ab versus the architecture being wrong, drove most of what happened after.

By the end of January there was HTTPS, sessions, a crypt module, and a mux router in progress. The tree got "refactored whole" twice in the first month. Normal early-framework churn. February was the first real mistake.

NukeDB

We spent two weeks of February 2026 building a database feature called NukeDB. The design: SQLite running on :memory: as the query front, mmap as storage, and a background thread syncing to disk. Writes go to memory, and no request ever waits on fsync.

It worked in a demo. Then the commits started: "prevent nuke db double free", "fix segfault when finding missing data", and a week more in that vein. A hand-rolled durability layer is a research project. We were trying to build a web framework. NukeDB was replaced by vendoring SQLite and using it directly. The design doc is still in the repo at blueprints/nuke_db.puml, as a warning.

The rule that came out of it: the custom component you find exciting is usually a distraction from the boring one that matters. We kept the boring ones, the router, the memory model, the reactor. The rest came off the shelf.

Benchmarks

In May a benchmark harness landed in CI. It runs wrk with 12 threads and 400 connections on a 4-vCPU AMD EPYC runner, against Axum, Gin, and a Spring Boot service tuned with JDK 25 Leyden AOT. On the September 5 run, CWIST did about 152k req/s at 17 MiB RSS. Axum did 155k at 16.6 MiB. Gin did 116k at 30 MiB. Spring did 77k at 1.31 GiB.

Our own README carries the caveat: the runner's CPU model changes between runs, and that moves the numbers more than most code changes do. We keep the harness because it catches regressions. The ranking itself doesn't mean much.

mimalloc: a negative result, kept

Issue #25 asked whether swapping in mimalloc would help. We ran a real CI A/B (PR #34). Everything got worse. RPS down 1.8%, P99.999 up 97%, RSS up 120%. The C1M prefork model spawns many short-lived processes that allocate very little, and mimalloc reserves a per-process arena up front. That is exactly wrong for this shape. A single mallopt(M_ARENA_MAX, 1) (PR #35) won instead, for free.

Results like this usually get deleted and forgotten. We wrote it into the issue and closed it there.

io_uring: built, ripped out, rebuilt smaller

The io_uring story is the longest one. In May 2026 we built a completion-based io_uring backend with O(1) free-slot management. The full async model. In August we ripped it out. The ring infrastructure moved into the reactor as a drop-in epoll replacement, and the request I/O path went back to synchronous. The commit message says it directly: "revert(http): roll back io_uring coalescing/state-machine perf regressions." The state machine that the completion model forced onto the HTTP parser cost more than the syscalls it saved.

After that, the wins came back incrementally. Batch submissions. An eventfd registered with the ring. Wake coalescing. One io_uring_enter and one recv per keep-alive request. On-demand ring paging, which cut RSS 19%. In September, PR #187 landed the RX-uring receive path: one IORING_OP_RECV SQE instead of POLL_ADD plus recv. It measured neutral, 342.4k versus 341.7k rps, and we kept it because it removes a class of races. SQPOLL was evaluated and closed negative. Kernel SQ-thread wakeup on 6.12 can stall for milliseconds.

It took five months and one full revert to settle on io_uring as a better epoll, rather than an end in itself.

The tail lives before dispatch

The most useful single measurement came from a per-event latency probe shipped in v3.6 (PR #186). We expected tail latency to live in the handlers. It does not. Queue delay at p99 was 10 milliseconds. Callback execution at p50 was 25 microseconds. A request waits far longer to be handed off than it takes to run.

That finding reorganized the v3.7 priorities. The work is on the queue, not on user code.

C1M and the July patch storm

The headline mode is C1M, one million concurrent connections per machine, in the prefork tradition. It started as a typo, "C100K", in May. Then came "use c1m mode by default." Version 2.4.2 shipped with the commit message "fix busy-wait fundamentals instead of disabling the mode", because the alternative was quietly turning the mode off, and we did not want to do that.

The week of v2.1.1 through v2.4.2, five days in July, was the bill for all this. JWT clock-skew fixes. HTTP/2 and HTTP/3 stream resilience. And a full revert of HTTP/3 to the v2.1 base after a bad merge, rebuilt and re-shipped within days. v2.5.0 arrived the next week as the first "app-ready" release. It also fixed a fork-after-thread bug that had been sitting there since the multiprocess rewrite.

HTTP/3, three times

HTTP/3 took three attempts. The first was in May, a two-day sprint that pulled in ngtcp2, nghttp3, and monocypher, on top of HPACK fixes from the project's first external contributor. It worked in the lab. The August hardening wave shows what the lab means in practice: 0-RTT replay guards, FIN on empty bodies, interop against Firefox's neqo stack, certificate verification on by default.

In September we deleted ngtcp2 entirely (PR #100) and pinned lsquic plus BoringSSL. Fewer moving parts, one upstream to track. The policy for when upstream won't merge what we need is now written down: WebTransport lives on a dev branch only, and the roadmap says v3.7 ships without it rather than pinning main to a topic branch. We had already made that mistake once.

The cache that silently didn't work

This is the finding I would put on a poster. ADR-0001 exists because a member of the public watched our TCP traffic and noticed that CWIST_ENDPOINT_FIXED, the documented RAM-caching endpoint mode, was not caching. Sixteen identical requests invoked the handler sixteen times.

The investigation found more than a performance bug. When the cache did hit, it bypassed routing and authorization. Cache keys ignored the query string. Cached blobs were whole-wire captures that replayed the first response's connection headers. The feature was shipped and documented, and it was both useless and unsafe.

The sentence from the ADR worth quoting: "The framework cannot infer that assertion from a function pointer, absence of credentials, or a previously successful response." A cache is a correctness boundary, not an optimization. The fix is a new opt-in mode, CWIST_ENDPOINT_PUBLIC_FIXED, storing representations instead of wire bytes, bounded at 64 KiB bodies, 256 entries, and 60-second TTL. We did not silently upgrade the old behavior, because silently strengthening the contract of deployed declarations is how frameworks betray their users.

The cache that does work as designed is called Big Dumb Reply. It learns slow responses at runtime and replays them with sendfile. Entries are lock-free, blobs are swapped atomically, and reclamation is epoch-based. One cache was an accident. The other is a component.

A garbage collector, in C

v3.5 shipped a full garbage collector. This is a deliberately unfashionable choice in C, so it came with a measured price: about 3% overhead with the GC dormant, and about 86% with it fully on. The reclamation core is epoch-based GC from libttak, our companion library, which also provides the token buckets and the lock-free io_queue. Connection reclamation will not free an object a live thread still uses. There are TLS-destructor sweeps for threads and an atexit sweep for the process.

The interesting call was how to intercept malloc. We rejected link-level --wrap, because it would rewrite allocation inside vendored BoringSSL and lsquic, code that never includes our headers and has no reason to expect a non-libc allocator. Interception is opt-in per translation unit with a one-line macro instead. The docs give the reason: users will write malloc, not cwist_alloc. Meet them where they are.

WebAssembly, in four phases

The v3.6 theme was WebAssembly, tracked as four phases under one issue.

Phase 1 put the SQLite-backed cwist_db into the WASM build. It also fixed a strange bug: a tree-wide clang-format pass had rewritten "=>" as "= >" inside EM_JS blocks, corrupting format strings. CI now refuses to reformat those files.

Phase 2 shipped the cwist-wasm npm package with a first-party JS wrapper.

Phase 3 did the design work. Streaming dispatch at the WASM boundary, and sessions for a runtime with no persistent instance. The model is to pin the signing secret and put the state in a signed client-side cookie, so instance lifetime stops mattering. The crypto is a bundled header-only SHA-256/HMAC, so there is no dependency to leak. Phase 3 also landed WASI: preview1 under wasmtime, then WASI 0.2 serving real HTTP over wasi:sockets.

Phase 4 is a complete example app, routing, validation, templates, database, and pinned-secret sessions, running behind a Service Worker.

One decision I am glad we made, because it is boring: libcwist_wasm.a grew 5.2x when SQLite entered the bundle, and we evaluated splitting the build so that db-less apps do not pay for it. The measurement said the linker already handles this. An app that never calls cwist_db_* links about 64.8 KB. We closed the issue with no split build. If it ever matters, the lever is SQLITE_OMIT_*, not packaging.

Since v3.6, WASI preview1 has been retired, the wasm32-wasip2 target was promoted to supported with a CI gate, and a WIT component pipeline is being tried as a possible eventual replacement for the Emscripten bundle.

Process

A few process choices did as much work as any algorithm.

Minors ship when a theme is complete, not on a calendar. v3.5 was Full GC, v3.6 was WASM. The roadmap states it plainly: long release intervals, each landing a small number of large, well-tested changes.

Sanitizers run in CI permanently. The founding UBSan bug report came from an r/C_Programming thread. PR #24, from a first-time contributor, made the gate async-safe.

Commit messages stay honest. "revert(http): roll back io_uring coalescing/state-machine perf regressions." "http3: revert to v2.1 base." The log reads like a lab notebook, which makes it usable as a debugging tool.

About 191 of the 1,200 commits are CI automation, benchmark report updates and dependency syncs. Humans still review every behavioral change.

Where it stands

v3.6 shipped on September 21, 2026. v3.7 is in flight: WASI edge deployment, HTTP/3 client connection-close correctness, and the queue-delay work that the latency probe pointed at. v4.0, codename "Steady Amber", is the next generational release and the first we will call production-stable.

Roughly 282 of the 1,200 commits are from one person. The project's first-ever PR was a migration-safety fix in March. The same contributor sent the HTTP/2 HPACK fixes two months later.

What I would tell someone starting a framework

Benchmark on day three, and keep the harness in CI. Do not trust the ranking.

The component you fantasize about, a database in our case, is a trap. The one nobody looks at, the dispatch queue, is where your tail latency lives.

Famous infrastructure, io_uring, mimalloc, SQPOLL, gets adopted with a measurement in one hand and a revert plan in the other.

A cache is a correctness boundary. Ship it like one, or someone will watch your wire traffic and find out you did not.

Write down negative results and close the issue. Your future self greps for them.

Version by theme. Long intervals, few large changes.

When upstream will not merge, the branch is the bug. Ship around it.

CWIST began as a claim that C web development is still alive. Eight months in, the honest version of that claim is narrower. C web development is alive if you are willing to revert, measure, write things down, and delete the parts you are proud of.

Developers

Members

Contributor

P.S

This blog is written in C using CWIST.

Related posts

Back
Report

Comments

No comments yet.