Download this discussion as a PDF
- Number
-
0003
- Title
-
Guidedoc: Odin conversion engine
- State
-
committed
- Type
-
Standards Track
- Authors
-
Vikrant Rathore
- Created
-
2026-09-29
- Updated
-
2026-10-11
- Discussion
-
Local review; no external discussion URL assigned
- Labels
-
architecture
Guidedoc: Odin conversion engine¶
Abstract¶
This proposal defines Guidedoc, a reusable Odin library with a zero-allocation core, a native multi-SoA AST, and direct HTML and Typst renderers. It reads reStructuredText, implemented independently from the Docutils specification, and CommonMark 0.31.2. Native Typst uses an explicit toolchain adapter, and PDF is always compiled from Typst. There is no writer-lowering layer and no speed claim. The companion Guidedog CLI discussion, GDS 0002, defines host services and publication. The engine is implemented; the evidence and open items are at the end.
Motivation¶
The existing zenfmt design supplies useful document and ownership ideas, but its Zig APIs and allocator strategy do not express the intended Odin library contract. docodin, an Odin engine that reproduces Docutils byte for byte, supplies measured research on handles, struct-of-arrays storage, and source spans. Guidedoc has a different goal: its own semantic HTML and Typst output from the language specification. This proposal specifies that contract and separates the allocation-free engine from resource acquisition and the Typst runtime.
Purpose and boundaries¶
A universal converter needs a common language for documents. It must retain meaning until the chosen target forces a decision. A heading remains a heading; a table remains a table; an unsupported feature becomes an explicit loss or a refusal. A successful conversion must never conceal that distinction.
“Universal” describes an extensible architecture, not a promise that every format is implemented or that every pair can preserve all information. The supported inputs are reStructuredText, CommonMark, and Typst markup. The output targets are HTML and PDF. PDF always goes through Typst: reStructuredText and CommonMark produce Typst source; native Typst input retains its own semantics. Office formats and Markdown output are outside this design.
Three responsibilities¶
The Odin core accepts reStructuredText or CommonMark bytes and returns HTML or intermediate Typst bytes with structured diagnostics. A separate Typst adapter compiles generated or native Typst source into HTML or PDF as appropriate. The core does not open paths, print messages, look at environment variables, fetch URLs, or discover plugins. The application decides which files may be read and which artifacts may be published. An editor, a server, and a command line client therefore share the same conversion semantics.
The host may allocate while acquiring input or preparing a workspace. Guidedoc’s conversion procedures may not allocate. The Typst compiler has its own memory management; the zero-allocation guarantee does not extend into that dependency.
What we retain, and from where¶
The sibling repositories were inspected at ../zenfmt and ../docodin. zenfmt’s
ZDS 0013 supersedes the tree and lowering sections of its ZDS 0002; its
core/src/host.zig, options.zig and report.zig also inform this proposal.
docodin’s DOD 4 (data-oriented storage), DOD 12 (parser research) and DOD 13 (source
spans) record measured Odin experience. These are design inputs, not ported code.
| Inherited idea | Odin decision |
|---|---|
| Flat semantic tree (zenfmt) | Typed IDs and native SoA slices over fixed storage. |
| Handles, not pointers (docodin) | distinct u32 node IDs; bounds checks stay on. |
| Hot columns, cold side tables | Kind and subtree end are scanned; payloads are cold. |
| Byte spans and line index | Spans are byte ranges; lines are computed on demand. |
| Writer lowering (zenfmt) | Omitted. Direct HTML and Typst renderers. |
| Docutils output fidelity (docodin) | Not a goal. The specification is the baseline. |
| Elm-style reports | Structured evidence and concrete next actions. |
| Compile-time bundles | Ordinary imported descriptors in explicit slice registries. |
| Arena growth | Caller storage with checked capacity; core never grows an arena. |
| Mandatory adjacent manifest | Opt-in host preservation artifact. |
We do not translate Zig’s comptime type factories or generated routing matrix. Odin’s package imports and procedure values are enough to connect a reader and a writer. A small indirect call per stage is acceptable; dispatch per character is not.
Terminology¶
A workspace is borrowed mutable storage for one active conversion. A document is a validated immutable view of a workspace snapshot. A reader constructs that view. A pass transforms it. A renderer encodes HTML or intermediate Typst markup. An attribute is optional keyed information attached to a node; it plays the role zenfmt called a facet. A resource is an included source, image, or other byte asset whose acquisition belongs to the host.
A render check finds unsupported target features before emission. It does not search for alternative representations. An artifact is a resulting byte stream. Publication means making a staged artifact visible at its destination. These terms distinguish conversion success from filesystem success.
Non-negotiable contracts¶
- No allocator calls inside core initialization, parsing, resolution, validation, rewriting, render checking, diagnostic construction, or bounded-buffer writing.
- No mutable global state. A registry may be shared when immutable; workspaces may not be shared by concurrent conversions.
- Every input-dependent loop has a bound or a progress invariant. Nesting has a
checked depth limit,
Limits.max_depth, whose excess is aLimitresult. The builder, validation, the HTML renderer, and the CommonMark reader keep their nesting in caller-provided stacks; see “Recursion is bounded by the depth limit and the stack budget” for the RST reader, the C++ declaration parser, and the other renderers. - Success returns valid output. Failure returns a stable code and a useful action. Corrupt input is a recoverable result, never an assertion or a panic. Content that does not fit a bounded buffer is a reported failure, never silently cut short.
- Declared losses are available before publication. Strictness can refuse them.
- Registry validation rejects duplicate format IDs; nothing is silently replaced.
convertvalidates its registry on every call, so an unchecked registry cannot convert with the first of two adapters. - Source maps use byte offsets into named UTF-8 sources. Display columns are a rendering concern, not a second coordinate system inside the engine.
Memory is part of the API¶
“Zero allocator” means that Guidedoc borrows already available storage and never
asks an allocator for more. It does not mean zero memory, zero copying, or an
unbounded document in a fixed buffer. Advancing a used-count inside a caller-owned
slice is allowed. Calling an allocator, including context.temp_allocator or an
arena allocator over a caller buffer, is not allowed inside the guaranteed boundary.
The host chooses stack arrays for small work, a fixed pool for a server, or one allocation per region for a CLI session. Large automatic arrays are unsuitable for small thread stacks. Documentation shows this choice without hiding it in a convenience constructor.
Storage regions and lifetime¶
| Region | Contents | Valid until |
|---|---|---|
| Input | Borrowed UTF-8 bytes and names, main and included | Last span or report use |
| Snapshot A / B | Node columns, payload tables, attributes, text pool | Reset of that bank |
| Scratch | Reader skeletons, stacks, symbol slots, render state | End of the stage |
| Reports | Fixed diagnostic records | Next conversion or reset |
| Output | Caller-owned byte buffer and used count | Caller reuses the buffer |
All regions must be disjoint, suitably aligned for their element types, and alive for the advertised lifetime. Typed slices avoid requiring callers to align raw bytes for the snapshot; scratch is the one raw byte region. Initialization checks lengths and rejects overlapping regions. Neither strings nor slices convey ownership in Odin, so borrowing rules appear in the procedure documentation.
Node text is always owned by the snapshot’s text pool. Readers copy text into the pool after escape processing and entity decoding. The input is referenced only through source spans, for diagnostics and source maps. This costs one linear copy and removes a whole class of lifetime errors: a snapshot never aliases input bytes or the other bank, so validation checks one pool and a rebuild copies nothing it cannot own.
Scratch is a marked bump region over a caller byte slice. A stage takes typed,
aligned slices with scratch_take, records a mark, and releases to it when finished.
Taking scratch advances a used count inside the caller’s slice; it is not an allocator
call. Exhaustion reports the scratch region and the minimum additional bytes.
Two snapshot banks permit rebuild passes. A pass reads A and writes B, validates B,
then swaps the active bank. On failure, A remains valid. An unchanged pass returns A
immediately. Bank B may have zero capacity when a caller runs no passes; a pass then
fails before it runs with Capacity and a need whose region is Second_Bank
(core.storage.second_bank), which names the cause rather than the nodes region. There is no persistent
sharing or copy-on-write mechanism in version one.
Sizing and carving cannot disagree about bank B. The one-buffer form,
workspace_size(n) and init_workspace_from_bytes(ws, slab), always plans one bank and
takes no bank argument. A workspace for passes is sized and carved from one value:
plan := slab_plan(n, second_bank = true), plan_bytes(plan), and
init_workspace_with_plan(ws, slab, plan). (An earlier draft gave both one-buffer
procedures a second_bank argument with a default; sizing with it and initializing
without it left bank B empty, and the failure surfaced only at pass time as a need for
one node.) init_workspace_with_plan is the authoritative path, which the one-buffer
form only chooses a plan for. All sizing
arithmetic is checked: slab_plan, plan_bytes, and workspace_size return
ok = false on overflow or a negative capacity, and initializing with such a plan is
Invalid_Input (core.storage.plan).
Every time a bank is cleared to be built again its revision rises, and a Document
records the generation, the bank, and the bank’s revision. Passes alternate banks
inside one generation, so the revision is what keeps a handle to bank A invalid after
a second pass has rebuilt bank A. Re-initializing a workspace advances its generation
too. The staged procedures also check that the document belongs to the workspace they
were given (core.document.foreign).
Small vocabulary, ordinary Odin¶
Use Type_Name for exported types and verb_noun for procedures. Call sites use a
short package alias, as in gd.convert. Borrow the clarity of raylib’s small verbs and
explicit lifecycle, while avoiding hidden global engine state. There are no fluent
builders, class hierarchies, required macros, or allocator defaults.
Inside this repository packages import each other by relative path, so a plain
odin build cmd/guidedog works. Embedders may vendor lib/ and map a collection.
import gd "lib/core"
import "lib/defaults"
// The host supplies typed backing slices for both snapshot banks, scratch bytes,
// source slots, diagnostics, and an output byte buffer.
ws: gd.Workspace
if r := gd.init_workspace(&ws, storage); r.status != .Ok do return
defer gd.reset_workspace(&ws)
result := gd.convert(&ws, defaults.registry(), gd.Request{
source = {name = "intro.md", text = "# Welcome\n"},
reader = "commonmark",
renderer = "html",
policy = gd.default_policy(),
limits = gd.default_limits(),
}, output[:])
if result.status == .Ok {
html := output[:result.written]
// Consume html before reusing the output array.
}
Initialization borrows storage; reset invalidates views and clears used-counts.
No close or destroy procedure is needed because the library owns no resources.
convert starts a new workspace generation and invalidates earlier workspace views,
even when the new call fails. It returns output length, report span, and a by-value
primary failure record. Callers needing inspection use the staged interface.
Status :: enum u8 { Ok, Invalid_Input, Unsupported, Policy, Capacity, Limit,
Cancelled, Internal }
Region :: enum u8 { None, Nodes, Text, Attributes, Lists, Links, Cells, Sources,
Scratch, Reports, Output, Source_Map }
Capacity_Need :: struct { region: Region, minimum: u64, exact: bool }
Result :: struct {
status: Status,
written: int,
reports: []Diagnostic,
primary: Diagnostic,
need: Capacity_Need,
}
init_workspace :: proc(ws: ^Workspace, storage: Storage) -> Result
reset_workspace :: proc(ws: ^Workspace)
convert :: proc(ws: ^Workspace, registry: Registry, request: Request,
output: []u8) -> Result
read_document :: proc(ws: ^Workspace, reader: Reader, request: Request) -> Read_Result
transform_document :: proc(ws: ^Workspace, doc: Document, passes: []Pass,
limits := Limits{}) -> Read_Result
check_render :: proc(ws: ^Workspace, doc: Document, renderer: Renderer,
request: Request) -> Result
render_document :: proc(ws: ^Workspace, doc: Document, renderer: Renderer,
request: Request, output: []u8) -> Result
Returned Document values carry the workspace, its generation, the bank, and the
bank’s revision. view and the staged procedures reject stale handles and handles of
another workspace. Render checking and emission use the same immutable snapshot,
renderer, and options; emission repeats the check. Applications must not mutate
exposed read-only-by-contract slices. Odin does not enforce this borrowing model.
Read_Result carries the same status, reports, primary failure, and capacity fields
as Result, plus a document valid only on success. Request includes limits, policy,
render options, optional passes, a resource provider, and a cancellation probe.
Defaults come from explicit default_limits and default_policy procedures;
zero-valued limits do not silently mean unlimited. default_request(name, text, reader, renderer) assembles a complete request from them.
A reader’s own settings travel in Read_Options.extension, a Read_Extension: the
caller’s pointer and its typeid, made by read_extension(&settings). A reader takes them
back with reader_settings(ctx, T) (or no_reader_settings(ctx) when it takes none),
which refuses a pointer of another type with Invalid_Input and a diagnostic naming both
types’ packages (core.read.settings_type, core.read.settings_none) before anything is
read. Settings meant for one reader are therefore never reinterpreted as another’s, which
a bare rawptr allowed: rst.Settings passed to the Markdown reader returned Ok and
read their memory as commonmark.Settings. The mechanism stays generic, so readers of
other registries use it with their own types. lib/defaults adds typed helpers for the
built-in readers (rst_options, sphinx_options, myst_options), which the compiler
checks as well.
Resources are provided, not fetched¶
An RST include names a file only when the parser reaches it, so a host cannot supply
every resource before parsing. Request.resources is a Resource_Provider: a procedure
value and caller-owned user data. The reader asks for a name relative to the including
source; the provider returns borrowed UTF-8 bytes or a refusal. The provider belongs to
the caller’s trust boundary and may allocate; the core never learns a path.
The core records each provided source in a caller-sized source table, so every span names its origin. Include depth, cycles, and total included bytes are bounded. A missing provider, a refused name, and a cycle have distinct diagnostics. The CLI provider confines reads to approved roots with symlink-aware checks, as GDS 0002 says.
Capacity is a result, not an exception¶
We do not promise to determine parse storage from input length alone. Nested
constructs, decoded text, tables, and extensions can expand differently. The host
provides capacities and an overall budget. A failing append reports the exhausted
region and the minimum needed at that point. This is a lower bound unless exact
is true. Arithmetic overflow is checked before addition, multiplication, or slicing.
A host may restart a pure conversion with more storage, subject to its budget. Retries start from input bytes; no partial parser state is resumable in version one. The CLI allows at most three retries and reports when the next size would exceed the chosen budget. That policy is outside the allocation-free library.
A renderer has one encoding procedure and three emitter modes: check reports
unsupported content without output, measure counts bytes, and write stores them.
The engine runs check and measure together, then write. An undersized output buffer
yields its exact byte requirement before modifying output. Parsing or render checking
failure also leaves output untouched. An unexpected failure or cancellation during
emission returns written = 0; callers treat all output bytes as invalid. A
successful write checks that actual bytes equal the measured count.
Callbacks and custom passes are part of the caller’s trust boundary. Core cannot guarantee that an arbitrary caller-supplied procedure performs no allocation. There is no streaming sink API in the first release; partial writes and backpressure deserve a separate contract.
Conversion sequence¶
Failure at steps 3 through 6 returns a diagnostic and prevents emission. After the caller consumes the result, reset or another conversion releases all borrowed workspace views together. Concurrent callers use separate workspaces.
Proving the allocation boundary¶
Tests install a failing allocator in both context.allocator and
context.temp_allocator, then exercise successful and failing conversions; any
allocator call fails the test. tests/allocation_test.odin converts one document
with every built-in reader (RST, Sphinx, MyST, strict CommonMark, and Sphinx Markdown)
into every renderer, across the options that choose other code (Math_Style.Plain and
.Sphinx, permalinks, a highlighter), with and without a rebuilding pass, and checks
that the output still holds every piece of content, including a heading, an admonition
title, and a citation label longer than any fixed buffer. A refused allocation that
code ignored would otherwise surface as quietly missing text, as it once did for
Sphinx-style display math.
Counting context allocators alone is insufficient evidence, so
tests/boundary_test.odin also audits the source of every package inside the boundary,
on tokens so comments and strings never count. The boundary is derived, not listed: the
test follows the repository imports of the packages whose code runs inside a conversion
(lib/defaults, which assembles the readers and renderers, and
lib/highlight/guidedoc, the highlighter renderers call) transitively, less declared
host packages, so core, the readers, the renderers, and lib/highlight are covered, and
a package a conversion comes to import is covered without editing the test. Each may import only an
allowlist of Odin packages (base:intrinsics, base:runtime, core:unicode,
core:unicode/utf8, core:slice, core:strings, and core:encoding/entity), and
from those only procedures that do not allocate, such as strings.index but not
strings.clone; it may not call make, new, append, delete, free, or the other
allocating built-ins, declare dynamic arrays or maps, or name the context’s allocators.
Test files are outside the boundary. Host and Typst costs are measured separately.
Allocation-free helpers that a host also wants in an allocating form keep the pure form
in core and the allocation in the host: normalize_nfd_into(dst, s) writes Unicode NFD
into caller storage (measuring with a nil dst), and the project host allocates for its
index keys.
Use Odin’s native #soa slices for the AST and typed side tables. Callers supply
backing arrays; initialization borrows their SoA slices. Never use #soa[dynamic].
Representation follows the work¶
reStructuredText and CommonMark readers share a semantic document representation. The HTML and Typst renderers traverse it directly. There is no writer-lowering layer, capability-cost graph, alternative search, or loss optimizer. Each renderer states exactly how each supported construct is emitted. A missing mapping is a reported implementation gap, never a silent fallback chosen by a planner.
Native Typst follows a separate path. Typst is both markup and a programming language. Extracting all its semantics into a small Odin tree would require an additional evaluator and still discard target-dependent behavior. Guidedoc instead passes native Typst to the Typst toolchain for HTML or PDF, through an explicit adapter outside its allocation-free core.
The conversion routes¶
HTML exported from Typst must be checked against the pinned toolchain’s supported features. The product does not promise that arbitrary page geometry has an exact HTML equivalent. A target limitation produces an actionable diagnostic, with a direction to adjust target-specific Typst content. Native Typst HTML uses the compiler’s export behavior; it is not a home-grown conversion of Typst tokens.
Flat structure, meaningful types¶
Use one preorder node forest with a synthetic document root. Node_Id is a distinct
u32; index zero is the root and the maximum value is reserved as invalid. Text is
an offset and length into the snapshot text pool. Source_Id names an entry in the
workspace source table, so spans in included files are distinguishable.
Node_Id :: distinct u32
Text :: struct { start, length: u32 }
Source_Span :: struct { source: Source_Id, start, end: u32 }
Node :: struct {
kind: Node_Kind, // hot: scanned by traversal
flags: Node_Flags,
subtree_end: u32, // hot: exclusive preorder end
value: i32, // small per-kind scalar: heading level, number
text: Text, // leaf content, URL, label, or name
payload: u32, // row in the kind's payload table, or NONE
span: Source_Span,
attributes: Attribute_Range,
}
nodes: #soa[]Node // a tag scan reads nodes.kind without loading payloads
A subtree rooted at i occupies [i, subtree_end[i]). The first child, if present,
is i + 1; the next sibling is the current child’s subtree end. Skipping a subtree is
constant time. The builder keeps an open-node stack and patches a node’s end when it
closes. Preorder does not mean that immediate children form a flat slice.
| Index | Tag | End | Meaning |
|---|---|---|---|
| 0 | Document | 5 | Contains the entire document |
| 1 | Paragraph | 5 | Contains inline content |
| 2 | Text | 3 | “Read ” |
| 3 | Strong | 5 | Nested emphasis |
| 4 | Text | 5 | “carefully” |
Rich payloads live in three typed SoA tables: lists (style, start, delimiter,
tightness), links (destination, title, reference name, resolved node, link kind), and
cells (row and column spans, alignment, header flag). Other per-kind data is either a
scalar value or a keyed attribute row. Attribute keys are an enum — class, id, name,
language, alt, width, height, scale, align, format, title, and a custom key carrying its
own text — so common keys never copy strings. Resource bytes are supplied by the host
and referred to by source ID; the core never fetches them.
Readers build skeletons in scratch and emit in preorder. CommonMark needs every link
reference definition before inline parsing, and a paragraph may later become a setext
heading or shrink when definitions are removed. Its reader therefore builds a block
skeleton in scratch, then emits final nodes with inline content in one ordered walk.
RST sections close when a later title of equal or higher level appears; the builder’s
open-node stack expresses that directly. No reader patches a closed subtree. A Sphinx
only directive may hold sections; the reader places it where its first title’s level
belongs before opening it, so this holds there too.
Semantic coverage¶
The block kernel is: document, section, heading, paragraph, block quote, attribution, bullet and enumerated list, list item, definition list, definition item, term, classifier, definition, field list, field, field name, field body, option list, option item, option group, option, option argument, option description, line block, line, code block, doctest block, math block, raw block, thematic break, table, table head, table body, row, cell, figure, caption, legend, image, admonition, topic, sidebar, rubric, container, compound, footnote, citation, target, comment, substitution definition, contents, and a generic directive container for third-party extensions.
The inline kernel is: text, emphasis, strong, code, link, reference, image, footnote reference, citation reference, substitution reference, inline target, inline math, raw inline, soft break, hard break, subscript, superscript, title reference, abbreviation, and a generic role span carrying its role name.
The catalog grows with conformance fixtures when a specified construct needs a semantic distinction. The generic directive container is for third-party extensions, not an excuse to label standard RST unsupported. Cross-document references use stable project keys, not local node IDs.
No hash table may grow implicitly. Symbol tables use caller-owned scratch slots with a bounded load factor and capacity errors. Collision and probe limits prevent hostile input from causing unbounded work.
Resolution is a stage, not a rebuild¶
After parsing, the reader resolves names into side tables: named, anonymous, and indirect hyperlink targets; auto-numbered and symbol footnotes; citations; and substitution definitions. A reference’s link row records its resolved node. A substitution reference points at its definition, and renderers emit the definition’s content in place. Nothing is copied or spliced, so resolution needs neither bank B nor a pass. Substitution cycles, duplicate targets, and unresolved references have dedicated reports with both source sites where two exist.
Reader conformance is explicit¶
reStructuredText means the language specification, implemented independently. The normative baseline is the Docutils reStructuredText markup specification with its standard directives and interpreted text roles. Guidedoc reproduces the language, not Docutils’ Python API, node classes, HTML writer, or message texts; docodin already provides that fidelity for projects that need it. Implementation phases do not redefine the baseline as a permanent subset. Release claims identify the tested specification revision and publish a coverage ledger for every construct, directive, and role, backed by Guidedoc’s own fixtures.
The ledger covers indentation, sections, transitions, all list families, literal and doctest blocks, line blocks, block quotes, grid and simple tables, explicit markup, comments, substitutions, targets, anonymous and indirect links, footnotes, citations, directives, roles, and escaping. Readers preserve source spans through substitution and reference resolution. Include cycles, substitution cycles, duplicate targets, unresolved references, and invalid nesting have dedicated reports.
Some directives require authority or a service: include, raw, images, math, and
external resources. Recognizing and validating their syntax is part of full support.
Executing them is governed by explicit policy. A disabled include is reported as a
policy refusal, not mislabeled a parse error or silently discarded. No shell or Python
runs because a document requests it. Sphinx-specific domains and custom directives are
an extension layer and do not belong to the base RST conformance claim.
Markdown means CommonMark 0.31.2. Run its complete example suite. GFM tables, footnotes, task lists, and strikethrough are not silently included in the base mode. Optional extensions require explicit names, separate fixtures, and documented interaction rules. HTML blocks and inline raw HTML are parsed according to CommonMark even when publication policy refuses them.
Typst means native Typst markup and its language semantics. Do not invent a “Typst-lite” parser. Compiler diagnostics are adapted into the same report model. Native source maps to itself; generated Typst carries an interval map from generated ranges to original RST or CommonMark spans.
Direct rendering rules¶
Each renderer implements one documented mapping per semantic construct. A section
becomes an HTML section with a heading, or a Typst heading; a footnote becomes an
HTML note with backlinks, or a Typst footnote. A table with spanning cells preserves
spans in both targets.
Every renderer routes node kinds with one exhaustive switch over Node_Kind: a
new kind does not compile until each renderer says which of its parts writes it, or
that it writes nothing (a comment, a substitution definition). A part given a kind it
was not written for reports core.render.kind (Internal) instead of dropping the
node and its subtree, and an integration test renders a document holding every kind
with every built-in renderer. Payload tables (payload_of), content models, and
translation units are decided the same way.
Math. RST and CommonMark math is LaTeX notation; Typst math is a different language.
HTML carries the LaTeX source in a math class for a client-side service. The Typst
renderer uses a bounded LaTeX-to-Typst math adapter for fractions, roots, scripts,
Greek letters, operators, delimiters, accents, matrices, and common symbols. An
unknown command is an explicit Unsupported diagnostic naming the command; policy
may instead emit the source verbatim with a warning. It is never silently dropped.
Raw HTML cannot be executed as Typst. The PDF path refuses target-specific raw content by default and points to a portable construct or a target-specific alternative. An explicit omission policy may warn and omit it; this is a fixed policy, not writer lowering. Standard semantic constructs remain required mappings.
HTML text, attribute values, and URLs have separate escaping rules. URL policy rejects
executable schemes. Raw HTML is refused by default unless trusted-content mode is
selected; the CLI exposes this as --raw=refuse|omit|allow. The CommonMark conformance
harness uses a reference-rendering profile so safety policy is not confused with
parser conformance.
Generated Typst uses a fixed trusted preamble and escaped data. Ordinary source text
never becomes executable Typst code through interpolation: every markup-significant
character is escaped, and code, URLs, labels, and resource paths have dedicated
string-literal encoders. A user’s native .typ file is executable Typst content
within the configured toolchain environment.
Validation and cost¶
Validation checks column capacity, used bounds, root coverage, nested and strictly advancing subtree ends, legal parent-child categories, payload rows and text spans, source IDs, table geometry, and resolved references. Each finished document and rebuild must pass validation before rendering. Assertions are for internal bugs; untrusted data always receives a recoverable diagnostic.
For n nodes, t text bytes, and r references, normal traversal and emission are
linear in n + t + output bytes. Reference indexing is expected linear within its
probe budget. No implementation may promise linear time for the entire RST grammar
without measurement. A work limit bounds repeated delimiter scans, reference
processing, and substitution expansion.
A rebuild uses two banks and is linear in copied nodes and text. Version one favors a few clear stages over persistent trees, SIMD, or speculative parallel parsing.
Recursion is bounded by the depth limit and the stack budget¶
Decision. The builder’s open nodes, validation, copying between banks, anchors, the
HTML renderer’s walk, and the CommonMark reader keep nesting in caller-provided stacks
taken from scratch. The RST reader (a body inside a directive, list, or quote parses its
own body), the Sphinx layer’s C++ declaration parser (templates, declarators, and
expressions nest), and the text and Typst renderers (a node renders its children) are
recursive descent instead. Every one of those recursions passes through a guard
(enter, cpp_enter) that checks two limits before the next frame:
Limits.max_depth, the nesting the content may have, whose excess iscore.limit.depth;Limits.stack_bytes, the machine stack the recursion may use, whose excess iscore.limit.stack. AStack_Guard(lib/core/stack.odin) records the address of a local where the reader or renderer starts, and each guard compares the address of a local in the current frame with it. The measure is exact whatever the procedures’ frame sizes, the compiler, or the optimization level.
Both are Limit results; no input and no max_depth can overflow the stack. The
default budget, DEFAULT_STACK_BYTES, is 512 KiB (0 means the default). Measured in
the default build, the deepest constructs take about 1.3 KiB per level of nesting, so
the default max_depth of 256 needs at most about 330 KiB and reaches its depth limit
first; a build without optimization has frames about 2.7 times larger and may stop at
the stack budget sooner, which is still a Limit. The thread running a conversion needs
its caller’s stack, plus stack_bytes, plus STACK_RESERVE (128 KiB) for the leaf work
below the deepest guard (inline parsing, the math adapter’s own bounded recursion). The
default fits a 1 MiB thread, the smallest default thread stack of the supported
platforms; Odin’s threads on Unix take RLIMIT_STACK, usually 8 MiB.
The C++ parser also bounds its steps by the length of the declaration (64 steps per
byte), since backtracking can retry alternatives at every level of nesting; a fold
expression is tried only when its ... is present, so nested parentheses parse in
linear time.
Alternatives. Rewriting the recursive readers and renderers around explicit
caller-provided stacks would remove the machine-stack dependence entirely, but touches
most of their procedures for no change in behavior once the guard exists. Estimating a
worst-case frame size per level and refusing a max_depth that cannot fit is fragile:
the frames change with every edit, compiler version, and optimization level, and the
estimate must hold for every path. Running conversions on threads sized from the limit
needs thread stack sizes that Odin’s core:thread does not expose; a host that wants
more nesting raises stack_bytes and runs the conversion on a larger thread itself.
This supersedes the earlier decision of this section, which bounded recursion by
max_depth alone and made a host that raised it responsible for a larger stack.
guidedog convert --stack-kib=N sets stack_bytes for one conversion, beside
--max-depth. guidedog build takes the same --stack-kib, --max-depth, and
--max-nodes, and every session of the build (reading, translating, loading a cached
document, linking and rendering a page) uses them; they are part of the digest of what
changes reading and output, so changing one reads and writes again. The engine’s limit
reports name the Limits field to raise (limits.stack_bytes, limits.max_depth,
limits.max_nodes), since the engine knows no command line; lib/cli replaces those
hints with the flags that set them (--stack-kib=N, --max-depth=N, --max-nodes=N)
when it shows the reports.
The largest budget is what the thread running the conversion holds besides the frames
around the guarded recursion: host.stack_budget_limit leaves 2 MiB (and at least half
the stack), so 6 MiB of Linux’s and macOS’s 8 MiB, and the 512 KiB default of Windows’
1 MiB. Convert runs on the main thread, whose stack is the soft RLIMIT_STACK on Unix
(host.main_stack_bytes). A build reads on the main thread, or with -j on threads
core:thread starts, which it sizes to the soft RLIMIT_STACK too, or leaves at the
system’s default (512 KiB on macOS) when that limit is unlimited
(host.thread_stack_bytes); pages render on the main thread. So a build’s cap is the
smaller of the two stacks’ limits, and a larger --stack-kib is refused (build.stack)
before anything is read. Without the option, the budget is the default, or less on a
thread too small for it (host.default_stack_budget). Windows gives every thread the
executable’s 1 MiB.
The supporting libraries¶
The libraries Guidedog is built from may not import lib/core (the independence test
lists what each may import), so each bounds its own recursion, with the same two ideas:
a limit the input’s nesting meets the same way on every platform, and, where the
recursion’s cost per level is not fixed, a measure of the machine stack itself. Each
reports a clear error or leaves the construct out with a warning; none crashes.
lib/jinja:Options.max_nesting(default 100) bounds the depth of the tree the parser builds, counting brackets, nested tags, and every link of operator, filter, test, attribute andelifchains, since chains build left-deep trees; every later walk over a template is bounded with it.Options.stack_bytes(default 512 KiB) is a frame-address guard started by the first public call on the environment and checked by the parser’s nesting and everyevaland statement, so macro calls, includes and recursive loops (max_recursion, 200) cannot overflow either. The defaults agree, so an ordinary template always meetsmax_nestingormax_recursionfirst and the guard is only the backstop for hostile ones: each procedure a level of nesting or recursion passes through keeps a small frame, dispatching through tables and leaving the rest of its work to helpers that are never inlined, so a level of recursion takes at most about 1.1 KiB in an optimized build and 1.9 KiB in a debug one (x86-64 Linux and arm64 macOS), and 200 levels fit in half and three quarters of the budget. The tests check both limits, and that margin, at every optimization level and under AddressSanitizer, whose builds get four times the budget. Walks over values (repr, equality, ordering,tojson) stop at 200 levels: repr prints[...]for a list that contains itself, as Python does.lib/treesitter: tree-sitter parses with a heap stack, so any nesting parses;deeper_thanmeasures a tree with a cursor, without recursion.lib/pyscan: a module whose syntax tree nests deeper thanMAX_SYNTAX_DEPTH(256) is not analysed and is markedtoo_deep(autodoc warnsautodoc.too_deep); Python itself refuses 200 nested brackets or 100 indentation levels. Name resolution was bounded per path byMAX_DEPTH(48), but paths multiply (modules that eachimport *from two others, functions decorated several times with the next), so one lookup also hasMAX_STEPS(100,000) steps andLOOKUP_STACK_BYTES(256 KiB) of stack, measured from its first frame, because an attribute of a class starts nested lookups through its MRO. MRO computations nest at mostMAX_MRO_DEPTH(100) classes deep; a longer inheritance chain gets an incomplete MRO instead of cubic time.lib/autodoc: annotation evaluation (MAX_EVAL_STEPS) and value description (MAX_DESCRIBE_STEPS) count their steps, since aliases and constants that each use the next several times multiply the work; past the budget the annotation or value is shown as written. One directive generates at mostMAX_GENERATED(10,000) directives, with the warningautodoc.limit, since a class that contains itself under its own name is documented again at every level.lib/odindoc: Odin’s parser (core:odin/parser) is recursive descent with no limit and takes 8 to 20 KiB of stack per bracket in a debug build, so an 8 MiB thread overflows near 600 nested parentheses, and odindoc cannot guard it from outside. Before parsing, it estimates the parser’s stack from the tokens (brackets, pending type prefixes, and cast, ternary andelsechains, each weighted by a measured frame cost) and leaves out a file pastPARSE_STACK_BYTES(1 MiB), with the warningodin.autodoc.nesting. Odin’s own core library estimates under 300 KiB. Thewhenconditions odindoc evaluates are followedMAX_WHEN_DEPTH(64) levels deep.lib/gettextalready bounded its plural expressions (100 levels, 128 nodes). Four costs a hostile catalog could square are now linear in it (or n log n): a PO field’s continuation lines are gathered in a builder and copied once, not concatenated line by line; an MO file whose strings share bytes is refused as damaged (sorting the table’s spans finds any), so a small file cannot make every entry check and copy the same large string; the fuzzy search ofmergewalks a length-sorted, trigram-ranked candidate list that stops where no candidate could win, giving the answer comparing every pair gives, within a work bound of 256 bytes read per byte of msgids (Merge_Statscounts it); and wrapping a long line counts only the next piece.lib/gettext/cost_test.odinchecks each by what it allocates or counts.lib/highlightkeeps nested lexer states in an explicit stack of 24 frames; andlib/graphdogdoes not recurse (Graphviz’s parser reports DOT nested 10,000 deep as an error).
Alternatives. One shared guard package would remove the small repetition of the
frame-address guard in jinja and pyscan, but would make every library depend on it;
the guard is a dozen lines. Running Odin’s parser on a thread sized from the file
would avoid the estimate, but Odin’s core:thread does not expose stack sizes.
Package structure and implementation order¶
Package boundaries express dependency direction; file boundaries keep procedures and explanations small enough to read in one sitting.
cmd/guidedog/ main.odin; composition and process exit only
lib/core/ package guidedoc; memory, AST, builder, validation, reports
lib/readers/commonmark/ CommonMark 0.31.2; myst-parser 0.16.1 syntax through a hook
lib/readers/rst/ RST blocks, inline markup, directives, roles, resolution
lib/readers/sphinx/ Sphinx directives, roles, domains, for RST and Markdown
lib/render/html/ semantic HTML encoding and escaping
lib/render/typst/ Typst source encoding, math adapter, source maps
lib/render/text/ plain text, as Sphinx's text builder writes it
lib/defaults/ explicit built-in reader and renderer registry
lib/host/ storage sizing, retries, resource provider, publication
lib/typst/ adapter to the Typst backend (installed CLI or the bridge)
lib/cli/ options, commands, terminal diagnostics
lib/gds/ metadata, lifecycle, registry, promotion, recovery
lib/project/ discovery, configuration, symbols, cache, builders, themes
lib/jinja/ Jinja templates for pages; independent of Guidedoc
lib/highlight/ code highlighting with Pygments' classes; independent
lib/highlight/guidedoc/ the highlighter as a Render_Options.highlighter
lib/graphdog/ Graphviz's C library for graphs; independent
lib/gettext/ gettext catalogs for translations; independent
cmd/graphdog/ a dot-compatible command over lib/graphdog
native/typst_bridge/ narrowly scoped Rust-to-C ABI bridge to the Typst compiler
native/graphviz/ script that builds Graphviz statically for lib/graphdog
tests/ conformance, malformed input, goldens, integration
tools/ source-limit checks and documentation commands
docs/gds/ standalone discussions, index, template, and archive
core imports no reader, renderer, OS, terminal, or Typst package. Readers and
renderers import core and nothing that allocates. Defaults imports those adapters and
assembles a registry. Host and CLI compose the pure library; Typst remains an explicit
host dependency. Embedders can import core and a selection of adapters without the CLI.
Guidedog, the application, is cmd/guidedog with lib/cli, lib/project, and
lib/host; Guidedoc, the engine, is lib/core with the readers and renderers. The
libraries a Guidedog build also uses (Jinja, the highlighter, graphdog for Graphviz, and
gettext) import no Guidedog or Guidedoc package, so other programs can use them alone;
only the small lib/highlight/guidedoc adapter joins one to the engine.
tests/independence_test.odin checks every one of these import rules, and
tests/examples_test.odin type-checks every program under examples/ and every complete
program in the cheat sheets (README.md, lib/**/README.md) and the guide (a Markdown
block fenced as odin, or an odin code block), so the documented examples cannot fall
out of date with the API. A block marked fragment is compiled too, completed by one
convention: its own import lines stay at file scope and the rest becomes the body of
main, so a fragment names its imports and declares the values the prose around it
describes. Examples in doc comments, an indented block after a line “Example:” (Odin’s
convention, which odindoc renders as a code block), are compiled the same way, and code
indented in a comment without that marker fails the test.
tests/api_docs_test.odin requires a doc comment on every procedure the libraries
export, which states its contract as GDS 0001 asks: what it borrows or copies, how long
results live, and how it fails. Memory has a checked shape, by each package’s
convention: a package that allocates nothing (core, the readers, the renderers,
defaults, highlight) may not call an allocator anywhere; in one whose procedures
allocate, a procedure whose own code allocates, frees, or takes an allocator names who
owns the result (the caller, an arena, the allocator, or what it borrows); and a package
that runs entirely in the caller’s arena (cli, gds, pyscan, autodoc, odindoc)
says so once, in its package comment or its README’s “Memory” paragraph. What only the
core uses is private.
The readers and renderers are narrowed the same way. Each exports its entry points
(VERSION, CONFORMANCE, the descriptor procedure, and the read, read-inline, or render
procedure), its settings types with their defaults, and, where another package builds on
it, a named interface: the RST reader’s extension interface (the directive and role
handlers, the parsed directive and its options, the parser’s node, scratch, line, and
report procedures, and the host interface through which MyST runs the same directives),
the inline parser’s hooks, the CommonMark reader’s extension interface, and the Sphinx
layer’s configuration, MyST settings, and autodoc host interface. Everything else is
private: a file that holds only implementation is #+private, and an implementation
declaration beside public ones carries @(private). Diagnostics are identified by their
stable code; the Message values behind them are private, except the one a dialect
reports itself (rst.DIRECTIVE_UNKNOWN). Package-internal tests still reach private
declarations; a test in another package uses the public surface. lib/README.md
(“Reader and renderer surfaces”) lists each surface.
Every other library is narrowed the same way, so each can be used on its own through a
small API that its README cheat sheet lists: cli exports only run; gds its record
workflow (load, plan, lock, apply, recover, build), with the path and string helpers it
once exported moved to their users; highlight its lexing, writing, style, and problem
procedures, with the lexer engine private, and highlight/guidedoc only the two
highlighters; treesitter and graphdog their procedures, with the C functions behind
them private; pyscan the analysis autodoc builds on, with the scanner and analyser
private; autodoc and odindoc their sessions or projects and one document
procedure; jinja, gettext, and i18n the API their cheat sheets document, with
tables, limits, messages, and the members behind procedure groups private; inventory
and fetch their load, find, write, and get procedures. Where the narrowing found an
unclear owner, the contract was fixed rather than described: fetch.Result is now
always the caller’s, with fetch.destroy, and i18n.figure_for_language always returns
a string the caller owns. typst exports its backends, locate, the two compiles, and offset_of, with the bridge’s C types private. lib/host and lib/project are narrowed by the same rule, below.
Host surface¶
lib/host is narrowed by the same rule. It exports what the command line, the project
builder, the engine’s test suites, and an embedder need to run the engine with memory
and files: sessions that size, grow, and retry storage within a budget (Session,
run, read, session_destroy, output_text, describe_all, budget_report), the
steps of a session for a caller that loads saved documents (plan_session,
grow_session, prepare_session, resize_output, resize_source_map, read_kept, read_owned),
storage sizing without a session (sizes_for, make_storage, destroy_storage),
budgets (Budget, budget_join, budget_leave, Ledger, charge, discharge,
settle, budget_problem, Meter, meter_allocator, metered_refusal), files
(read_source, read_file, same_contents, each_chunk, io_problem), includes and
confinement (Files, files_provider, canonical_roots, confined, confinement,
containing_root), publication (publish; write_unsynced, sync_all, folders_of
for a commit of many files), diagnostics (Report, Problem, describe, exit_code,
display_path, the text and JSON renderers, excerpt_at), interrupts, stack budgets,
and the guidedog convert route (convert_artifact). Budget reservations, region
growth, staging files, terminal escaping, and byte counts as text are private.
One ownership rule covers the package: what a procedure returns is allocated in
context.allocator and belongs to the caller, which frees it with that allocator’s
arena; arguments are borrowed for the call. The exceptions say so in their contracts: a
session owns what its results refer to until it runs again or is destroyed, and a
resource provider or meter borrows its state for as long as it is used. Failure is a
Problem or false, never a panic; a budget refusal is host.budget, a system
refusal within the budget host.memory. lib/host/README.md lists the surface;
tests/api_docs_test.odin requires a contract comment on each procedure in it, and
tests/surface_test.odin requires each exported name to appear in the cheat sheet, so
the surface cannot grow without being advertised. The project builder’s surface is
recorded with the project (the 0006-sphinx-replacement draft, “Project surface”).
Sessions remember the allocator that supplied their current regions, so changing the
calling context cannot change who frees them. Project steps and CLI conversions use
reclaiming allocators for those regions: retries release actual storage rather than
only returning a budget charge while an invocation arena retains the memory. Buffer
replacement reserves the complete new buffer while the old one remains live, then
frees the old buffer before returning its charge. The same budget and policy govern
initial preparation and subsequent growth. read_owned inputs belong to the session,
including their ownership records, and are freed before their reservations are released.
read_kept preserves its original contract: the caller owns and frees the input bytes,
and the session releases their charge when it is destroyed. Cached objects use
read_owned, so discarding a stale cache frees its real storage immediately.
Diagnostic coordinates count Unicode scalars in the original source. Terminal excerpts are bounded windows around the error span, with separate cell widths for escaped controls, combining marks and wide characters. JSON retains the complete original source line, so clipping or escaping cannot change an editor’s coordinates.
odin run tools/check enforces the source limits: lines of at most 99 characters, never
more than 108, files of at most 1,408 lines, and, for Odin, procedures of at most 70
logical lines. The line and file limits cover all implementation source, including the
CSS, JavaScript, HTML templates, Typst templates (such as the quickstart book template),
shell scripts, Rust, and TOML shipped with the code. Documentation is exempt: Markdown,
reStructuredText, and Typst prose under a docs directory (these records included) are
written for readers, and third-party data under data directories is
kept as published.
odin build cmd/guidedog -out:build/guidedog
odin test lib/core -out:build/core-tests
odin test tests -out:build/guidedog-tests
The Odin toolchain is dev-2026-09:a2fb372b7; Typst is 0.15.1.
Adapter contracts¶
A registry is an immutable pair of slices: reader descriptors and renderer
descriptors. Stable IDs identify adapters: rst, sphinx, commonmark, myst, html,
typst (the generated-source renderer), and text. A reader ID names one profile, never
a default that settings switch: commonmark is pure CommonMark 0.31.2 and takes no
settings, and myst is the MyST Markdown that projects write (their .md documents are
read with it and the Sphinx layer’s directives and roles). The CLI maps md to commonmark
and exposes pdf
as a host route through the Typst renderer and backend. Native .typ input is a
backend route, not a falsely advertised zero-allocation reader.
A descriptor holds ID, version, conformance status, and a procedure value. Options are
typed: Read_Options carries the reader’s own settings as a typed Read_Extension, and
Render_Options
selects standalone or fragment output, the reference profile, and a document title.
Registry validation checks unique IDs, non-nil procedures, and required metadata;
convert runs it before reading, and the staged procedures refuse a reader, renderer,
or pass without a procedure. Adapters retain no workspace pointers between calls.
Core’s exported names form two surfaces, and everything else is private to the package:
the conversion surface (workspaces, slabs, requests and their defaults, convert, the
staged procedures, results, diagnostics, and the read-only view), and the adapter and
pass surface (descriptors, contexts, the builder, the emitter, scratch, stacks, buffers,
symbol tables, anchors, and shared text helpers), with three services for hosts beside
them (documents saved as bytes for a build cache, Unicode NFD, and line and column
positions). Generation and bank bookkeeping, slab carving, the pipeline’s helpers, the
kind and payload tables, the fold tables, the object format’s constants, and the
diagnostics only core reports are internal. Procedures on the two public surfaces state
what they borrow, how long results live, and how they fail; lib/README.md lists every
exported name by surface, and a test fails when core exports a name the list lacks. A Buffer whose content does not fit sets a sticky overflow flag and
counts the bytes it was asked for, so a caller measures into a zero buffer and then
takes exactly that much scratch (scratch_buffer) instead of truncating into a fixed
array.
A pass is a procedure plus caller-owned user data. It receives an immutable source
snapshot and a destination builder; it returns unchanged, rebuilt, or failure. The
engine runs passes in order and validates each rebuilt snapshot. The builder exposes
begin_node, end_node, add_leaf, add_text, add_attribute, and typed payload
procedures. Every procedure returns a checked result. A failed builder is poisoned
until reset, so ignoring one capacity failure cannot publish a truncated document.
Unbalanced nodes fail finalization.
Storage holds two Snapshot_Storage values, scratch bytes, source slots, and a
diagnostic slice. Each snapshot contains native SoA node, attribute, list, link, and
cell slices and a text pool. All lengths are capacities, with separate used-counts.
Tables may have zero capacity when a profile does not use them; encountering such a
construct yields a precise capacity result. No descriptor constructor allocates, and
no hidden “default workspace” exists.
The host library offers convert_artifact with explicit source, target, resource
provider, workspace sizing, staging destination, and optional Typst backend. Its
allocation boundary differs from gd.convert: core conversion is zero-allocation;
resource acquisition and Typst compilation may allocate. A caller requesting PDF
without a backend gets an actionable result.
Diagnostics without allocation¶
Diagnostic :: struct {
code: string, // stable, static: "rst.reference.unresolved"
severity: Severity,
status: Status,
span: Source_Span,
title: string, // static
message: string, // static template: "No target is named {0}."
hint: string, // static direction
args: [2]Diagnostic_Arg, // integer, source span, or static text
}
Messages are static templates with at most two typed arguments, so constructing a
report never formats text. A source-span argument is quoted by the host renderer from
the source table. When report storage fills, primary keeps the first failure, a
counter records suppressed reports, and a truncation flag is set. Capacity failures
remain reportable because they need no storage beyond the by-value primary record.
Correctness before tuning¶
Native SoA storage, bounded loops, and explicit lifetimes are foundational choices. They do not imply a speed claim. First make the complete pipeline correct, safe, and understandable. Prefer a clear copy to a fragile alias. Bound CPU work and memory at all stages; report exhaustion before overflowing counters or indexing storage.
After end-to-end correctness is established, profile representative manuals and adversarial inputs. Record peak storage, high-water counts by region, CPU time, and output size. Optimize only demonstrated bottlenecks. SIMD, parallel parsing, and cache-specific tuning are deferred. No throughput target is a release requirement.
Acceptance matrix¶
| Concern | Evidence required |
|---|---|
| Memory boundary |
|
| AST safety |
|
| CommonMark |
|
| RST |
|
| HTML / Typst |
|
| Native Typst |
|
| Diagnostics |
|
| Host transactions |
|
| Resource bounds |
|
Unit tests target semantics and invariants. Golden fixtures show deliberate output changes. Expected failures are results, never crashes. A PDF text check cannot replace visual review, and a screenshot cannot prove reference resolution.
Security considerations¶
The engine processes untrusted markup. Its normative safeguards are checked arithmetic and slicing, bounded parsing and expansion, generation-checked handles, validated references, and target-specific escaping. Host authority is supplied only through the resource provider. No core procedure reads paths or executes code. Native Typst requires the separate host containment contract in GDS 0002.
Backwards compatibility¶
There is no earlier Guidedoc release. The proposal does not promise compatibility with zenfmt’s Zig ABI, manifests, or serialized AST, nor with docodin’s Docutils-compatible node classes or wire format. Changes to a published node schema require a separate GDS and an explicit compatibility decision.
Alternatives considered¶
A growing arena was rejected because it violates the allocation boundary. A pointer tree complicates fixed storage and ownership. Native SoA retains a clear semantic model while exposing compact columns. A single forest with typed side tables simplifies source coordinates and mixed block/inline nesting; separate forests remain a possible measured redesign. Writer lowering was removed at the user’s direction.
Reusing docodin’s RST engine was considered. It is complete and measured, but it reproduces Docutils’ document model and output byte for byte and allocates from growing arenas. Guidedoc needs its own semantic tree under a fixed-storage contract, so it implements the specification directly and treats docodin as research. Borrowing input spans for node text was rejected in favour of an owned pool; mixed ownership made bank rebuilds and validation harder than the linear copy it saved. A custom Typst evaluator would duplicate the toolchain and is not proposed.
Reference implementation and open questions¶
The packages listed above implement this proposal. lib/README.md is the embedder’s cheat
sheet and lib/core/AST.md the normative node mapping; examples/ holds runnable programs,
one of which converts with static arrays under a panicking allocator. Beyond the staged
interface, three raylib-style entry points make the common case one buffer and one call:
workspace_size, init_workspace_from_bytes, and convert_text.
Evidence on 29 September 2026, on macOS arm64 and Linux x86_64:
| Concern | Evidence |
|---|---|
| Memory boundary |
|
| Capacity |
|
| AST safety |
|
| CommonMark |
|
| RST |
|
| HTML / Typst |
|
| Math | 76 LaTeX translations compile; unknown commands are refused by name. |
The RST ledgers now record multiline substitution names and generated Unicode
punctuation support as implemented, with tests. The reader descriptor remains
conservatively Partial; corpus agreement is not a claim of complete Docutils
implementation or arbitrary extension compatibility. List_Info.delimiter is one byte, so
the bullets •, ‣ and ⁃ travel in a bullet attribute. Storage defaults are estimates
refined by capacity retries, not measured profiles.
Review history¶
29 September 2026: initial Odin proposal. Review removed writer lowering, selected full RST and CommonMark, selected native Odin SoA, deferred optimization, and restored the standalone Guidedog Discussion format and inherited lifecycle.
29 September 2026, implementation review: state that RST is implemented from the specification rather than ported from docodin; make node text owned by the snapshot pool; define scratch as a marked bump region; replace up-front resource entries with a provider; resolve references in side tables so bank B is optional; require readers to emit in preorder from scratch skeletons; define one encoder with check, measure, and write modes; add the LaTeX-to-Typst math adapter; use relative imports in-repository; replace the options schema with typed options; and specify static diagnostic records.
29 September 2026, implementation: record the implemented packages and evidence, the
one-slab entry points, the target-name attribute for embedded aliases, and project
links and anchor prefixes for multi-document books.
29 September 2026, engine review: remove the last allocations inside the boundary
(Sphinx display math, NFD normalization) and prove the boundary with an option matrix
and a token audit; report content that does not fit a bounded buffer instead of cutting
it; size and carve slabs from one checked plan; add bank revisions and workspace checks
to document handles; validate registries in convert; make core internals private and
name the two public surfaces; add default_request and typed read options; type-check
the examples; extend the source limits to all implementation source, with documentation
exempt; and record that RST and the renderers recurse under the depth limit.
30 September 2026, review: bound recursion by a machine-stack budget
(Limits.stack_bytes) as well as the depth limit, in the RST reader, the C++
declaration parser, and the renderers, so no input or max_depth overflows the stack;
bound the C++ parser’s backtracking; and route node kinds in every renderer through
exhaustive switches. Later the same day: bound the supporting libraries the same way
(Jinja nesting and stack, Python and Odin source nesting, the multiplying work of name
resolution and autodoc expansion), make their closed dispatches exhaustive, and add
convert --stack-kib. Later again: narrow every library to the API its cheat sheet
documents, list core’s exports by surface and test the list, check the memory side of
every public contract by package convention, and compile every Odin snippet in the
documentation, fragments and doc-comment examples included.
30 September 2026, API review: narrow the readers and renderers as the core was
narrowed. Their implementation is private (#+private files, or @(private) beside
public declarations); each keeps its entry points, settings, and the interface other
packages build on, documents every exported procedure’s contract, and the doc-comment
test covers them. The Sphinx layer uses its own ASCII text helpers instead of the RST
reader’s.
30 September 2026, host API review: narrow lib/host to the services its callers use
(“Host surface”), state the package’s ownership rule, give every exported procedure a
contract, and list the surface in lib/host/README.md, which a test holds the exports
to.
Sources and design provenance¶
The following sources distinguish language facts from decisions made by this proposal. Web sources were checked on 29 September 2026.
- Neighboring
../zenfmt/docs/zds/records/0002-zenfmt-architecture.typand0013-layered-document-ir.typ: AST, facets, reports, and host separation. Writer lowering from that design is explicitly excluded here. - Neighboring
../zenfmt/core/src/host.zig,options.zig, andreport.zig: inspected implementation contracts, not code to port verbatim. - Neighboring
../docodin/docs/dod/0004-data-oriented-storage.rst,0012-parser-research-0-2-0.rst, and0013-source-spans.rst: handles, SoA storage, and span rules measured in Odin. - Odin language overview: native SoA, slices, distinct types, procedures, and implicit context.
- RST specification and the accompanying standard directives and standard roles.
- CommonMark 0.31.2: the selected Markdown baseline and conformance examples.
- Typst language syntax and HTML export.
- Typst compiler crate 0.15.1: Rust compiler stages and resource environment; the proposed C ABI is ours.
Lifecycle event
2026-09-29: prediscussion → discussion. Assigned a permanent number and opened for discussion.
Lifecycle event
2026-10-11: discussion → accepted. Maintainer-requested implementation-status
reconciliation; current supported contract reviewed and deferred scope stated
explicitly.
Lifecycle event
2026-10-11: accepted → committed. Current supported design implemented; beta
review fixes and limitations are recorded. Native library and integration
suites with leak checks; sanitizer and embedded-backend validation;
docs/manual/evidence/beta-review-20261011.md. Deferred features are not
claimed as implemented.