Bez: Generating a browser engine from specs and tests

Source: tangled.org
72 points by nerdypepper 3 hours ago on hackernews | 20 comments

6

Configure Feed

Select the types of activity you want to include in your feed.

This repository has no description

6

Configure Feed

Select the types of activity you want to include in your feed.

1.2k 21 50

Clone this repository

https://tangled.org/burrito.space/bez https://tangled.org/did:plc:jwnvvtd4hsk4hvgtyu66uwfy

git@tangled.org:burrito.space/bez git@tangled.org:did:plc:jwnvvtd4hsk4hvgtyu66uwfy

For self-hosted knots, clone URLs may differ based on your setup.

Download tar.gz Download .zip

Commits 1.2k

Limits (two Dagu slots per lane, writers per job, gateway in-flight per model, one
landing at a time), nine conflicts confirmed from the scripts, areas with no conflict,
and one next change per conflict.

README.md

Bez - a generated web engine#

Generate a web rendering engine from specs and tests.

Why#

Building a web engine by hand costs hundreds of engineers and many years, so only a few companies can have one, and only they decide how the web works. Bez generates the engine instead: from the specifications, with the three shipping browsers (checked against each other) and WPT as the tests. Once the pipeline exists, each extra engine costs very little.

Goals#

  • Complete, or tree-shaken. The default build is a complete engine. For uses other than a web browser, a build can instead contain only the web features its content uses: analyse a site or an app, and every feature it never touches is left out of the binary, not just switched off (docs/content-scoped-engine.md).
  • Fast on any device. A smaller engine does less work and needs less memory.
  • Easy to embed. A small engine with a clear API, for apps and devices that today ship all of Chromium or go without.
  • Fastest to update. When a spec or test changes, the affected code is regenerated and re-verified, not rewritten by hand.

How it works#

flowchart LR
  spec["spec text"] --> model["model writes<br/>many candidates"]
  model --> engine["candidate runs<br/>inside the engine"]
  browsers["Chromium, Firefox,<br/>WebKit (cached)"] --> check{"same boxes,<br/>same places?"}
  engine --> check
  check -- no --> model
  check -- yes --> code["committed as<br/>ordinary Rust"]

Status#

Coverage history: share of browser-compat-data leaf keys per status, one bar per recorded run, with a line behind the bars for the number of leaf keys reached out of the total

Computed against browser-compat-data 8.0.4 (17259 leaf keys, 74 rules in docs/feature-map.json) on 2026-09-25.

area generated hand-written linked oracle only unreached
overall 0.6% 0.3% 0.5% 5.7% 93.0%
css 2.5% 1.1% 2.0% 0.0% 94.4%
html 0.0% 0.0% 0.0% 0.0% 100.0%
api 0.0% 0.0% 0.0% 9.9% 90.1%
javascript 0.0% 0.0% 0.0% 0.0% 100.0%
svg 0.0% 0.0% 0.0% 0.0% 100.0%
webassembly 0.0% 0.0% 0.0% 0.0% 100.0%
http 0.0% 0.0% 0.0% 0.0% 100.0%
mathml 0.0% 0.0% 0.0% 0.0% 100.0%

See docs/dashboard.md for what these statuses mean and how this table is kept in sync.

DOM, style, box tree and fragment tree are hand-written in crates/dom and crates/layout and pass their browser checks (roadmap.md → "Phase 1 — Engine bootstrap"). Nine CSS 2.1 layout rules live in crates/layout/src/generated; eight were written by models and admitted against the three-browser vote, and block height keeps its hand-written rule because no model candidate beat it. Together they pass all 227 recipe cases and all 11 usable WPT normal-flow pages. Which rules the engine actually calls: roadmap.md → "Which resolver each generated function runs". What would prove the thesis: roadmap.md → "Proof of concept".

Findings so far#

  • Three browsers mostly agree, and the third names the odd one out. 699 of 705 browser-pair comparisons agreed across 235 documents (commit 3ca7323); all six disagreements were twelve-deep percentage nesting, and majority voting named Firefox the outlier every time. docs/firefox-app-units.md.
  • That Firefox difference is a real web-compat bug. Gecko rounds lengths to 1/60 px where Blink and WebKit use 1/64 px, which makes a flex item wrap only in Firefox or offsetWidth differ by a pixel. Reproductions match Mozilla-diagnosed breakage on Slack, Google Store and Samsung, and Mozilla bug 1719314. experiments/firefox-compat/.
  • The WPT vote table agrees across most of the suite. A stable three-engine majority covers 2,162,676 of 2,282,301 test/subtest keys (94.8%). docs/wpt-votes.md.
  • Script-observable behaviour agrees too. Trace probes under a fake-media profile agree on 15 of 18 engine pairs; the three disagreements are a real platform difference. docs/dump-format.md → "Trace protocol".
  • Conformance suites are usable oracles. Counting a test usable when two of three engines agree: WPT canvas 82.6% (92.7% excluding tentative), Khronos WebGL with dEQP 99.7%, WPT Web Audio 74.4% (85.6% excluding tentative); the WebGPU CTS about 85% for validation and about 50% for numeric execution. docs/conformance-oracles.md, docs/webgpu-cts.md.
  • Logic programming earns a narrow place. Margin collapsing written as Datalog agreed with the browsers on 1195 of 1195 offsets and caught a seeded bug the geometry check missed in all 227 cases. docs/logic-layer.md.
  • The economics hold for a subset of the platform. About 55–60% of engine-relevant compat entries have a usable automated oracle and generatable spec prose; roughly 8–18% have none. docs/platform-economics.md, docs/oracle-coverage.md.
  • An engine scoped to one site's content is designed, and its content analyser has landed. docs/content-scoped-engine.md.

Open questions#

What building each rule costs by hand, how wide a feature's generated part should be, how much of WPT is reachable without JavaScript, and where IDL-generated code ends: roadmap.md → "Open questions". Area-by-area status: docs/platform-areas.md.

Sources#