RFD0020 - Krasny Pretty Printer
- Feature Name:
krasny_pretty_printer - Start Date:
2026-03-23 - Status:
implemented
Summary
Section titled “Summary”This RFD proposes a new standalone package, krasny, as Riot’s owned OCaml
formatter.
krasny should:
- accept
Syn.Parser.parse_resultbut require a successfulSyn.build_cstbefore formatting - lower typed
Syn.Cstplus concrete token access into a formatting document IR - solve documents at a fixed width of
100columns - render formatted output through an
Std.IO-friendly writer API - expose one formatting style for everyone, with no user-facing knobs
The formatter should be the only rendering system Riot builds for formatted OCaml output. We should not build a separate fix-only fragment printer and then later build a real formatter beside it.
Motivation
Section titled “Motivation”Riot is heading toward two adjacent capabilities:
- syntax-directed lint rewrites
- a canonical
riot fmt
When synthetic rewrites need fresh syntax materialization, Riot will need some way to turn structured syntax back into OCaml source. Building one “small renderer for fixes” and later a second full formatter would be a bad split. Both systems would grow to cover similar syntax and layout concerns, and the repo would end up paying for two rendering pipelines.
The cleaner path is to start the formatter now and let all future rendering needs grow from that single subsystem.
Guide-level explanation
Section titled “Guide-level explanation”Contributors should think of krasny as a syntax-to-document pipeline:
Syn.parse_resultSyn.build_cst- typed CST plus concrete token access
- formatting document IR
- layout solver
- printer
Std.IOwriter
At a high level:
synremains responsible for parsing and faithful syntax structurekrasnyis responsible for layout and rendering policyriot fmtbecomes the user-facing entrypoint tokrasny
One formatter, one style
Section titled “One formatter, one style”krasny should be explicitly opinionated:
- no per-project knobs
- no style variants
- no local escape hatches beyond what OCaml syntax itself requires
If the style changes, Riot changes it globally and deliberately.
Primary path: CST-driven formatting
Section titled “Primary path: CST-driven formatting”Formatting should primarily traverse Syn.Cst, not raw Ceibo nodes and not
raw source text. CST gives the formatter the structure it actually wants:
- expressions
- patterns
- types
- declarations
- module items
- function parameters
- records and variants
That makes the formatter easier to reason about than a raw token/node printer.
Lowering should be structurally driven:
- pattern match on typed
Syn.Cstconstructors for semantic shape - read concrete keywords, punctuation, and delimiters from structured token access
- build
Doc - let the solver and printer decide spaces, line breaks, and indentation
That means lowering should not:
- slice raw source text
- check
String.starts_with "let rec "or similar - emit indentation or hardcoded whitespace strings like
" = "as a layout shortcut
If the current public CST surface does not expose enough token structure for a
construct, Riot should improve syn rather than teaching krasny to guess
from text.
No raw-text fallback
Section titled “No raw-text fallback”krasny should only format from a successful CST lift.
If Syn.build_cst fails, formatting should fail with a structured error. Riot
should not add a Ceibo-to-doc or raw-source fallback path for normal
formatting.
Reference-level explanation
Section titled “Reference-level explanation”1. Package boundary
Section titled “1. Package boundary”Introduce a new package:
packages/krasnyResponsibilities:
- document IR
- CST-to-doc lowering
- layout solver
- printer
- rendering to
Std.IO
Non-responsibilities:
- parsing
- linting
- rewrite planning
2. Public API shape
Section titled “2. Public API shape”The formatter should expose an Std.IO-friendly writer surface with one fixed
style and one fixed line width.
A first-pass API could look like:
module Krasny : sig type format_error = | Cannot_build_cst of Syn.build_cst_error
val format : Syn.Parser.parse_result -> (string, format_error) result
val write : writer:Std.IO.writer -> Syn.Parser.parse_result -> (unit, [> `Format of format_error | `Write of Std.IO.write_error ]) resultendEven if the first implementation renders into a string internally, the API
should be shaped around writing because that is how riot fmt and future
callers will naturally use it.
3. Internal pipeline
Section titled “3. Internal pipeline”The formatter pipeline should be:
Syn.parse_result -> Syn.build_cst -> lower -> doc -> solve(width = 100) -> print -> writerThere should not be a second fallback formatting pipeline that reconstructs syntax from raw text.
4. CST and token boundary
Section titled “4. CST and token boundary”Syn.Cst should provide the semantic structure that lowering wants, but the
formatter still needs access to concrete tokens such as:
letrec=fun->withstruct/sig/end
That access should come from structured CST token fields or from deliberate token access on the attached red subtree, not from source slicing.
For example, a let binding should expose semantic information such as:
is_recursive- pattern
- parameters
- value
and should also expose the concrete tokens that matter for rendering, such as:
let_keywordrec_keyword : token optionequal
If a CST node does not expose enough token structure for faithful lowering, the
fix belongs in syn, not in krasny.
5. Document IR
Section titled “5. Document IR”krasny should not render directly by string concatenation. It should lower
syntax into a document IR.
A first-pass document model should be in the Wadler/Oppen family, for example:
module Doc : sig type t = | Empty | Text of string | Space | Line | Hard_line | Concat of t list | Indent of int * t | Group of t | Flat_alt of { when_flat : t; when_broken : t; }endThe exact constructor list is less important than the architectural point:
- CST decides structure
- token access decides concrete syntax atoms
Docdecides layout- the solver decides breaks and indentation
- the printer renders the solved document
Lowering should treat spacing and breaking as document structure, not as string assembly. In particular:
- use explicit doc items for spaces and breaks
- keep tokens and whitespace separate
- do not embed layout into strings like
" = "or"\n " - do not consider indentation while lowering beyond building
Indent/Groupstructure into the document
6. Layout engine
Section titled “6. Layout engine”The layout engine should follow the broad shape used by modern document-based pretty-printers:
- a document tree that carries grouping and break opportunities
- a
fits-style solver that chooses flat or broken layouts under a width constraint - a printer that renders the solved document
The implementation does not need to be academically pure. It does need to be:
- deterministic
- simple to reason about
- able to grow without rewriting the architecture later
- fixed at
100columns unless Riot deliberately changes the global style
7. Comments and trivia
Section titled “7. Comments and trivia”Comments and trivia must be handled intentionally.
Because the formatter is CST-first, but comments live naturally in Ceibo
trivia, krasny will need a comment attachment strategy when lowering syntax
to documents.
This RFD does not prescribe the full comment algorithm yet, but it does make one call explicit:
- comments are part of formatting design, not a post-processing hack
8. Why this is not a fragment printer
Section titled “8. Why this is not a fragment printer”This RFD explicitly rejects building a fix-only “render one node to text” printer as a separate subsystem.
If Riot later needs to materialize synthetic rewrites, that rendering should
come from krasny’s document pipeline, not from a second rendering codepath
inside riot-fix.
That means:
- fragment rendering may eventually exist
- but it should be a mode of the formatter, not a separate tool
Drawbacks
Section titled “Drawbacks”This is a substantial new subsystem.
Even a simple formatter will require:
- broad syntax coverage
- comment handling
- document/layout infrastructure
- long-term stability discipline
It also introduces pressure to define formatting behavior for syntax that Riot does not yet lint or transform, because a formatter must cover the whole language surface.
Rationale and alternatives
Section titled “Rationale and alternatives”Alternative: build a small fix-only renderer first
Section titled “Alternative: build a small fix-only renderer first”Rejected.
That would create a second rendering system that overlaps heavily with the formatter we already know Riot will want.
Alternative: print directly from Ceibo without CST
Section titled “Alternative: print directly from Ceibo without CST”Rejected as the primary design.
Ceibo is the right source of truth for lossless syntax, spans, and trivia, but formatting wants structured syntax more than raw red-tree traversal. CST is the better primary traversal surface, with deliberate token access where concrete syntax matters.
Alternative: direct AST/CST printer with no document IR
Section titled “Alternative: direct AST/CST printer with no document IR”Rejected.
OCaml layout is rich enough that direct string printing would quickly collapse into handwritten spacing and line-breaking heuristics scattered across many node printers. A document IR gives the formatter a cleaner long-term shape.
Alternative: raw-source or Ceibo fallback when CST lift fails
Section titled “Alternative: raw-source or Ceibo fallback when CST lift fails”Rejected.
That would push formatter complexity into source slicing and syntax
reconstruction right where Riot most wants strong structure. If CST lift fails,
the correct fix is to extend syn’s faithful typed CST and its token exposure,
not to add a second formatting pipeline.
Alternative: wait until synthetic rewrites force the issue
Section titled “Alternative: wait until synthetic rewrites force the issue”Rejected.
Riot already wants riot fmt, and waiting would only encourage ad hoc
rendering logic to leak into unrelated systems first.
Prior art
Section titled “Prior art”Three formatter families are especially relevant.
Swift swift-format
Section titled “Swift swift-format”Swift’s official formatter is the closest structural model for krasny’s CST
boundary:
- typed syntax tree traversal
- concrete token access from the syntax tree itself
- insertion of formatting controls before and after real syntax tokens
- a dedicated pretty-printing stream and printer
The key lesson is not that Riot should copy Swift’s printer token-for-token, but that it should follow the same boundary:
- semantic structure from typed syntax nodes
- concrete spelling from explicit token access
- layout decisions in a dedicated pretty-printing layer
That is the clearest precedent for using Syn.Cst plus structured token access
instead of raw source reconstruction.
Source:
JavaScript prettier
Section titled “JavaScript prettier”Prettier is the clearest precedent for Riot’s document model and solver:
- a first-class
DocIR with explicit groups, indents, and line variants - a
fits-style solver - a separate printer over the solved document
That is the strongest reference for the doc -> solver -> printer half of
krasny.
Sources:
Rust rustfmt
Section titled “Rust rustfmt”rustfmt is useful mostly as a warning:
- it is explicitly heuristic rather than algorithmic
- many
rewrite_*functions returnStringorOption<String> - it uses source spans and snippets pervasively
That is exactly the style Riot should avoid in krasny. It is valuable as a
large-scale practical formatter, but it is the wrong starting architecture for
the cleaner typed-CST-plus-Doc design Riot wants.
Source:
Document algebras
Section titled “Document algebras”The underlying pretty-printing lineage from Wadler, Oppen, and later work is still the right conceptual base:
- build docs, not strings
- separate structure from layout
- make width-sensitive decisions late
That is the family of ideas krasny should build on.
Unresolved questions
Section titled “Unresolved questions”- What is the smallest good first
Docalgebra for OCaml in Riot? - How should comments be attached from Ceibo trivia to CST-driven formatting nodes?
- Which concrete tokens should be exposed directly on
Syn.Cstnodes, and which should remain reachable through generic token access on the attached red subtree? - How should future fragment rendering for synthetic fixes reuse the same pipeline without exposing formatting internals everywhere?