Stegstr Leaderboard

FINAL FIELD — frozen 2026-09-01. Every runnable entry is compared on how well hidden data survives lossy recompression, then independently app-tested for build, connectivity, reliability and the Nostr layer. Community evidence adds the rest. The winner has been picked.
Winner

#108 by @wordsmithakif

After a month of testing, this is the one. It's the most complete app that actually works: signed builds for Mac, Windows and Linux, a real CLI and MCP server, and nine upstream bugs fixed, including one that let a relay forge posts. It came through two rounds of heavy interaction testing without a single crash, kept its relay connections the whole time, and its robust JPEG mode survived every compression profile with no hint about the channel.

It's not perfect. Cancel doesn't cancel the identity conversion, delete has no confirm step, and profile picture uploads fail silently because nostr.build changed their rules. Those came from upstream and evrey fork has them. Honorable mentions: #71 for the resize-proof codec and the relay fixes (it just came in after the freeze), #83 for the cleanest architecture, and #79 for being the only one who fixed the upload problem.

Next step is combining the best parts of these into one app. Thanks to everyone who entered. — David

Evaluating? Start here. The instructions explain the gauntlet, how to get each version (public vs. on-request), the criteria, and how to report what you find. Then submit evidence — a bug, a strength, a screen recording — for any finalist. Every submission is tied to a signed-in account.
⬇ Evaluator instructions (PDF) Submit evidence →
18
Finalists (17 frozen + 1 post-freeze)
Still in contention
Knocked out
12
Independently app-tested
survives lost manual test

How survival is measured

01

Objective gauntlet

For entries that expose a command line, we embed a known message, push the image through 5 simulated lossy compression profiles locally, and decode. Same test for all — real, reproducible numbers.

02

Control check

Each engine must first recover its own uncompressed output. If it can't, we flag it rather than post a misleading score.

03

Human testing

Desktop-app entries (no command line) are tested by people actually using them, and their survival + reliability come from submitted evidence.

04

Independent app testing

Entries that build into a runnable app are exercised end-to-end on a clean Linux testbench: compile from the pinned commit, connect to a local relay, fuzz the UI with bursts of adversarial events, drop the connection to test reconnection, and verify the Nostr layer signs and publishes correctly. Findings are marked ◎ app-tested.

05

You decide

Operation, reliability, and security are weighed from these results plus community evidence. No single auto-computed winner — the evidence informs the final call.

Field frozen 2026-09-01. 119 entries were received; 40 active entries were triaged, entrants without a runnable deliverable were given outreach and time to supply one, and 18 finalists are everyone who did. One of those (#117) supplied its source only as an archive that could not be run for evaluation and has been removed pending a runnable copy. #71 was submitted after the freeze and added as an explicit, disclosed exception (marked "post-freeze add"); it is scored by the same process so it can be compared, not to change the frozen field.
Updated 2026-09-06. Added the Sep 5 real-channel (harness WhatsApp / Telegram profile) results for the desktop apps, the #79 Telegram-target retest, the interaction bugs the fuzzer found in #108, a three-cover gauntlet row for every scored entry (headline survival stays the frozen scoring-cover number), and community report counts from the evidence form.
Final-choice scope, 2026-09-07. The winner is chosen among entries that are social clients or messaging tools built on the codec. #103 is set aside on that criterion: its codec scored 100% blind and its relay publish is correct, but it has no feed interactions, replies, reactions, follows or DMs to compare against the others. #101 is out on its non-functional Nostr layer. That leaves #108, #71 and #83.
Heavy interaction battery, 2026-09-06 (top desktop pair + #83). #108 and #71 were built from their pinned commits and run through an identical battery on the Linux testbench: two seeded UI-fuzz sessions of 150 steps across four instances, two adversarial protocol swarms of about 305 actions each (reply chains 20 deep, delete-then-reply races, exact replays, future timestamps, floods, oversized and malformed tags, deletions of other users' notes), XSS and flood injection, and a relay drop. Neither crashed or logged an error. Both carry the same two upstream interaction bugs. They differ on connection resilience (see each row). #83, which is a messaging tool rather than a social client, got a protocol-level swarm instead and showed a delivery-blocking weakness in its default beacon mode. Every run is seeded and reproducible.
Hinted pass added 2026-09-06 — informational only. The gauntlet is blind by design: the encoder is never told which platform class it will face, because in real use the sender does not know either, and an entry that survives without a hint is superior to one that needs it. Ranking and knockouts use the blind result only. Entries that expose a platform or robustness option were additionally rerun with that option set per profile (dashed second row where present) so the difference is visible. Outcome: #107 reaches 60% when told the platform but 20% blind, so its knockout stands; #63 reaches 20% in its robust mode but 0% blind; #93 is unchanged either way; #96, #87 and #70 have no option to hint. Three desktop entries that were only "manual" were also scored through their own command-line tools (#108, #118) or MCP encoder (#79) by the same process. #59's command line turned out to be a placeholder that writes nothing, and #82's decoder recovers bytes but its CLI cannot delimit them — both knockouts confirmed. Separately, the Sep 3 phone calibration is withdrawn: every sample came back at 1024px on both platforms and in both WhatsApp quality modes, which is a transfer artefact, so results that depended on it are marked withdrawn rather than counted.
Resize vs. recompression. The real-channel (WhatsApp/Telegram) failures on the fixed-grid entries are driven by image resize misaligning the 8×8 block grid — not by recompression, which those codecs survive. Confirmed by porting an entry end-to-end, and consistent with the shared architecture; a resolution-independent design (#71) is the one that survives resize.
Survival numbers are real measurements from running each entry's own code; a “manual test” chip means the entry has no command line to script, so its survival is judged from human testing, not a fabricated number.
App-testing findings (◎) come from building each entry from its pinned commit and exercising it on a clean Linux testbench against a local relay. They describe behaviour observed in that harness; an entrant who believes a result is environment-specific can respond through the evidence form.
Compliance. The five compression profiles are internal simulations run on our own images. No third-party service is named or recommended for routing content.
Evidence is private to the contest owner; submitters sign in so activity is attributable.