Input format: Markdown, HTML - or something else entirely?
Markdown counts as settled, HTML is up for debate, and nobody has measured it for your stack. This lab announces no answer: you throw in your own material, ask your own question and decide blind which answer is better. Tokens and accuracy are two axes - counting alone measures the wrong thing.
01Your material, four versions
Paste your own content - Markdown or HTML, the lab recognises both. Four versions with the same content come out of it. The token count appears immediately, without a model call and without signing in. Counting uses cl100k (GPT-3.5/4); other tokenizers arrive at different numbers, but at the same ratios.
874 of 12000 characters
The numbers apply to the document with the sliders set below.
# Maintenance of the Nordfeld treatment plant Since the 2023 extension, the Nordfeld plant has run three aeration tanks instead of two. The change raised throughput, but also the power draw of the blower station. ## Operating hours In normal operation the blowers run 18 hours a day. In summer the demand drops, because degradation runs faster at higher temperatures; the controller then lowers the runtime on its own, down to 13 hours. ## The incident in March On 14 March blower 2 failed because of a defective frequency converter. The fault went unnoticed for eleven hours, because the alarm was sent to a phone number that had been discontinued. Since then alerting runs over two separate paths, and the call list is checked every quarter. ## Costs Replacing the frequency converter cost 4,180 euros. Retrofitting the second alerting path added another 900 euros.
02Your question, blind head to head
Two versions run in parallel against the same model. Which side is which format is not shown - someone who knows that Markdown is on the left will find Markdown better, and we would have a survey about prejudice instead of a measurement. The reveal comes only after your verdict.
This stack runs exactly one model, so there is no control here - both sides run against the same one anyway. Small models are more format-sensitive than large ones (Qwen3-8B varies by 9.18 percentage points across four formats); a result from here is therefore not a result for a frontier model.
The position control is not decoration: models find a fact at the start and at the end more reliably than in the middle ("Lost in the Middle", Liu et al. 2024, a drop of 20 percentage points and more). Measuring without it means measuring exactly that effect and calling it a format effect.
What has come in so far
What is shown is not "Markdown wins" but tokens per comparison won: the input tokens a format has cost across all its comparisons, divided by the comparisons it won. Lower is better. Confusing token savings with accuracy is the most common fallacy in this question - this number puts both axes together.
Still too few runs for an evaluation: 0 of at least 20 decided comparisons. A number from a handful of verdicts looks different tomorrow - so we only show it once it can take the weight.