Historical: Typsastra Khmer Render Preparation Implementation Plan¶
Retired: This document records a removed experimental design. Typsastra no longer rewrites Khmer preview/export input or inserts word-break controls. With
khmer_segmenterv0.3.0, the dependency is restricted to typing suggestions and spellcheck; its layout-oriented APIs are reserved for native layout-engine integration and are not called by Typsastra. The details below are retained only as historical context.
The original goal was:
User source files stay clean.
Generated preview/export files receive Khmer word-boundary markers.
Typst/Tinymist renders the generated files.
Editor, diagnostics, and reverse sync map back to original files.
Retirement status (2026-09-09)¶
The experimental render-preparation pipeline, setting, source directive, generated ZWSP mapping, and KHYP data have been removed. Typsastra now sends ordinary source to Typst, apart from unrelated preview-cache features such as Draft Preview. Future Khmer line-breaking work must be integrated natively with the layout engine rather than by an editor-side source transformation.
The implemented settings shape is:
The implemented source directive is scope-aware:
It disables render preparation for the following syntactic scope so examples can compare plain Typst rendering against Typsastra's generated ZWS rendering without changing application settings.
0. Core principle¶
Never modify the user’s .typ source files automatically.
Instead, Typsastra should maintain two document representations:
Authoring source
main.typ
chapters/intro.typ
Rendering source
.typsastra/cache/render/main.typ
.typsastra/cache/render/chapters/intro.typ
The rendering source may contain inserted U+200B zero-width spaces, but the authoring source should remain unchanged.
Phase 1: Minimal non-destructive renderer¶
Objective¶
Prove that Typsastra can compile a generated .typ file instead of the original source file.
At this stage, do not worry about reverse sync, diagnostics mapping, or perfect Typst parsing.
Output structure¶
For a project like this:
Typsastra should generate:
Source files are copied or symlinked into the render cache. Typst compiles:
instead of:
Recommended approach¶
Use project mirroring first.
That means the cache directory mirrors the original project structure. This avoids complicated path rewriting.
Preferred:
Source file → generated segmented copy
Asset file → symlink or copied file
Directory → mirrored directory
Example:
main.typ → .typsastra/cache/render/main.typ
chapters/intro.typ → .typsastra/cache/render/chapters/intro.typ
figures/diagram.png → .typsastra/cache/render/figures/diagram.png
refs.yml → .typsastra/cache/render/refs.yml
This keeps imports working:
#import "template.typ": *
#include "chapters/intro.typ"
#image("figures/diagram.png")
#bibliography("refs.yml")
because the generated render tree has the same relative paths.
Implementation tasks¶
Create a Rust module:
Initial public API:
pub struct RenderPrepareOptions {
pub enable_khmer_zws: bool,
pub project_root: PathBuf,
pub entry_file: PathBuf,
pub cache_root: PathBuf,
pub generate_source_map: bool,
}
pub struct RenderPrepareResult {
pub generated_entry_file: PathBuf,
pub changed_files: Vec<PathBuf>,
pub warnings: Vec<RenderPrepareWarning>,
}
Main function:
pub fn prepare_render_project(
options: RenderPrepareOptions,
) -> anyhow::Result<RenderPrepareResult>
Phase 2: Conservative Typst text scanner¶
Objective¶
Insert Khmer word boundaries only in safe visible text regions.
Do not try to fully parse Typst yet. Build a conservative scanner that avoids dangerous regions.
First supported segmentation region¶
Segment plain markup text:
Generated:
or preferably no visible space, only ZWS:
Skip these regions in MVP¶
Do not segment inside:
#import "..."
#include "..."
#let variable = ...
#show ...
#set ...
#bibliography("...")
#image("...")
#cite(<...>)
$ math $
`raw`
Also skip:
Scanner states¶
Implement a simple state machine:
enum TypstScanState {
MarkupText,
CodeExpression,
String,
Math,
RawInline,
RawBlock,
LineComment,
BlockComment,
}
For the first version, only transform text when:
Everything else is copied unchanged.
Conservative rule¶
When unsure, do not segment.
This is important. A false negative only means the preview is not perfectly broken in one place. A false positive can break the Typst document.
Phase 3: Khmer run detection and ZWS insertion¶
Objective¶
Within safe markup text, detect Khmer text runs, segment them, and insert ZWS between words.
Pipeline¶
safe text chunk
→ split into Khmer and non-Khmer runs
→ segment Khmer runs
→ insert U+200B between segmented Khmer words
→ preserve original punctuation and Latin text
Example input:
Possible generated output:
Do not insert ZWS¶
Avoid inserting ZWS:
before punctuation
after opening punctuation
inside existing ZWS-separated text
inside numbers
inside Latin words
inside URLs
inside Typst syntax
Suggested function¶
Return both text and mapping information:
Even if source maps are not used immediately, generate mapping data from the beginning.
Phase 4: Source map generation¶
Objective¶
Every generated render file should have a map back to its original file.
Generated files contain inserted characters that do not exist in the original source. Reverse sync and diagnostics require a map.
Source map file structure¶
For each generated file:
Example:
{
"version": 1,
"source_file": "/project/chapters/intro.typ",
"generated_file": "/project/.typsastra/cache/render/chapters/intro.typ",
"mappings": [
{
"generated_start": 0,
"generated_end": 12,
"source_start": 0,
"source_end": 12,
"kind": "original"
},
{
"generated_start": 12,
"generated_end": 15,
"source_start": 12,
"source_end": 12,
"kind": "inserted_zws"
}
]
}
Mapping kinds¶
For inserted ZWS:
That is a zero-length source span at the insertion boundary.
Required lookup functions¶
pub fn generated_to_source(
generated_file: &Path,
generated_offset: usize,
) -> Option<SourcePosition>
pub fn source_to_generated(
source_file: &Path,
source_offset: usize,
) -> Option<GeneratedPosition>
For reverse sync, the first one is more important.
Phase 5: Compile/export using generated entry file¶
Objective¶
Typsastra preview/export should compile the generated file.
Instead of:
use:
For Tinymist, the preview root should point to:
not the original main.typ.
UX behavior¶
User opens:
Typsastra internally previews:
The user should not need to know this unless debugging.
In dev builds, show a small status indicator when the experimental path is enabled:
Optional debug command:
Phase 6: Live preview update pipeline¶
Objective¶
When the user edits source files, regenerate affected render files and refresh preview.
Watch flow¶
User edits source
→ debounce
→ prepare affected file
→ update render cache
→ notify preview server
→ preview refreshes
Recommended debounce:
Important optimization¶
Do not regenerate the whole project on every keystroke.
Start simple, but design for this:
Changed file only → regenerate changed render file
Entry/import graph changed → regenerate affected files
Asset changed → update symlink/copy only
Cache key¶
For each .typ file:
cache_key = hash(source_content)
+ hash(segmenter_dictionary_version)
+ render_prepare_version
+ options
If unchanged, skip regeneration.
Phase 7: Diagnostics mapping¶
Objective¶
Typst diagnostics will refer to generated cache files. Typsastra should show them on original files.
Example Typst diagnostic:
Typsastra should map it back to:
because inserted ZWS may shift offsets.
Implementation¶
When receiving a diagnostic:
generated file + generated range
→ load corresponding source map
→ map generated range to source range
→ display diagnostic in editor
For inserted ZWS-only locations, map to nearest original position.
Rule¶
If mapping fails, show the diagnostic but indicate it came from the generated render file.
Do not hide errors.
Phase 8: Reverse sync¶
Objective¶
When the user clicks in preview, Typsastra should jump to the correct original source location.
Preview position comes from generated file:
Map it to:
using the source map.
Inserted ZWS behavior¶
If click lands on inserted ZWS:
map to the boundary between the two original words.
Suggested policy:
Inserted ZWS after word → source offset at end of previous word
Inserted ZWS before word → source offset at start of next word
Pick one and keep it consistent. I recommend mapping to the end of the previous word.
Phase 9: Better Typst syntax support¶
Objective¶
Gradually expand where segmentation is allowed.
After MVP, add support for visible strings in known functions.
Safe candidates:
Risky candidates:
Suggested strategy¶
Maintain a whitelist of visible-text contexts.
Example:
Only segment:
Skip:
Phase 10: Settings and controls¶
Required settings¶
Implemented application setting:
The UI row is hidden outside dev builds. The setting remains in the JSON schema so explicit local experiments are still possible.
Per-file/per-region directives¶
Implemented scope-aware comment:
Possible future region syntax:
This is useful for special cases, poems, code examples, or documents where exact spacing matters.
Phase 11: Testing plan¶
Unit tests¶
Test scanner behavior:
plain Khmer text → segmented
math → unchanged
raw block → unchanged
import path → unchanged
image path → unchanged
bibliography path → unchanged
comment → unchanged
mixed Khmer-English → only Khmer segmented
existing ZWS → not duplicated
Source map tests¶
Test:
generated_to_source on original text
generated_to_source on inserted ZWS
source_to_generated around inserted boundaries
diagnostic range mapping
multi-byte Khmer UTF-8 offsets
UTF-16 editor offsets if needed
Integration tests¶
Test project mirroring:
main.typ importing template.typ
main.typ including chapters/intro.typ
image path still works
bibliography path still works
generated entry compiles
Visual regression tests¶
Create PDFs/screenshots for:
narrow Khmer paragraph
justified Khmer paragraph
mixed Khmer-English paragraph
Khmer heading
Khmer figure caption
Khmer table caption
imported Khmer chapter
raw/code/math heavy document
Compare before/after render.
Recommended implementation order¶
Milestone 1: Render cache MVP¶
Deliverable:
Tasks:
1. Create render_prepare module.
2. Mirror project directory.
3. For .typ files, copy unchanged.
4. For assets, symlink/copy.
5. Compile generated entry.
No segmentation yet.
Milestone 2: Plain Khmer text segmentation¶
Deliverable:
Tasks:
1. Add conservative scanner.
2. Segment only MarkupText state.
3. Insert U+200B.
4. Skip math/raw/code/comments.
5. Add unit tests.
Milestone 3: Source map foundation¶
Deliverable:
Tasks:
1. Emit mappings while writing generated file.
2. Add generated_to_source lookup.
3. Add source_to_generated lookup later if needed.
4. Test inserted ZWS mapping.
Milestone 4: Preview integration¶
Deliverable:
Tasks:
1. Change preview entry path to cache/render/main.typ.
2. Watch original source files.
3. Regenerate changed render files.
4. Refresh preview.
5. Add UI indicator.
Milestone 5: Diagnostics mapping¶
Deliverable:
Tasks:
1. Intercept diagnostics from preview/compiler.
2. Convert generated path/range to source path/range.
3. Display mapped diagnostics.
4. Fallback gracefully if mapping fails.
Milestone 6: Reverse sync¶
Deliverable:
Tasks:
1. Capture preview-generated source location.
2. Map generated location to original source.
3. Jump editor to original file/offset.
4. Handle inserted ZWS locations.
Milestone 7: Visible string support¶
Deliverable:
Tasks:
1. Track simple function contexts.
2. Whitelist visible-text arguments.
3. Keep file paths/cite keys untouched.
4. Add tests.
Important technical recommendation¶
Start with ZWS only.
Do not insert SHY for Khmer in the initial render pipeline.
Use:
for Khmer word-boundary opportunities.
Leave:
for later Latin/technical hyphenation experiments, not Khmer segmentation.
This keeps the first version simpler and typographically safer.
Final architecture¶
The final system should look like this:
Original source files
↓
Typsastra file watcher
↓
Render preparation pipeline
↓
.typsastra/cache/render/*.typ
.typsastra/cache/maps/*.map.json
↓
Typst/Tinymist preview/export
↓
Diagnostics/reverse sync
↓
Mapped back to original source files
The user experience should be simple:
Write clean Khmer source.
Use normal Typst justification/tracking limits by default.
Enable experimental render preparation only when comparing generated ZWS boundaries.
Preview/export generated from the render cache can expose additional Khmer break opportunities.
No invisible characters are inserted into the source unless explicitly requested.
That is the current experimental promise. Do not present this feature as a production-default Khmer layout solution until segmentation quality is proven across real documents.