# oq public JS module API — v0 (experimental)

> **This package.** The module is [`src/public-api.js`](./src/public-api.js);
> this file is the contract. Snapshot lock: [`test/public-api.test.js`](./test/public-api.test.js)
> (`EXPECTED_SURFACE`). Types: [`src/public-api.d.ts`](./src/public-api.d.ts).
> Living sandhi spec: [`docs/sandhi-taxonomy-status.md`](./docs/sandhi-taxonomy-status.md).
> Package data flow: [`docs/DATA_FLOW.md`](./docs/DATA_FLOW.md).
> This package is the morphology source of truth.
> Engine bugfixes land here first; oq consumes a release after oq#888.

**There is no HTTP/REST API and nothing here to `curl`.** `oq` is a static
browser PWA with no backend server of its own — the closest thing to an "API"
is this JS module, `src/public-api.js`, which you `import` directly into
your own JS/TS code (browser or Node) and call as functions. The only things
you can actually point `curl` at are the third-party upstream data sources
`oq` fetches JSON from — see [`src/catalog/upstream-sources.js`](./src/catalog/upstream-sources.js)
for the URLs and [`docs/SOURCES.md`](./SOURCES.md) for each one's provenance,
licensing, and attribution requirements — those aren't an API `oq` provides,
just data it consumes. See
[Poking at upstream data directly](#poking-at-upstream-data-directly) below
for `curl`/`jq`/`sqlite3` examples against those sources.

Status of the surface defined by [`src/public-api.js`](./src/public-api.js)
(oq#394). Current version: **0.4.0**.

Version 0.4.0 moves source modules under `src/engine`, `src/catalog`, and
`src/generation`. `morph_engine.js` and `morph_analyze.js` are
`engine/morph-engine.js` and `engine/morph-analyze.js`. Deep imports on the
published archive follow that layout
(`/api/v0.4-latest/engine/morph-engine.js`). `v0.3-latest` stays the frozen
0.3 line and is not rewritten. Package subpaths `oq-api/engine`,
`oq-api/catalog`, and `oq-api/generation` export the direct modules; the
package root still clones shared results. `setActiveLocale` and
`getActiveLocale` are removed: pass `options.locale` or the third argument
of `t()`, and an omitted locale is English. `searchEntries` takes an options
object only; a string in that position throws.

Version 0.3.59 adds `generationCardFinishedState` / `projection_kind` so
Misiliineq finished status gates on honesty (partial, open obligations,
approximate, pending evidence) rather than `card.ok` alone, and so
recorded-chain projections stay labeled distinct from live generation.
Integration review: [`docs/reports/translation-integration-review.md`](./docs/reports/translation-integration-review.md).
Epic #280 stays open while oq-grammarian#828 / language review remain incomplete.

Version 0.3.58 projects generation candidates into mirrored sentence cards
(`projectGenerationCard`, modal EN→KL/DA→KL example views) without routing
retained sequences through `presentAnalysis`. Details are in
[`docs/generation-card.md`](./docs/generation-card.md). PWA state stays out of
oq-api; oq.gl consumes the documented integration contract.

Version 0.3.57 adds the public provider-injected sentence generation API
(`generateFromMeaning`, `generateFromSource`) that orchestrates interpretation,
grounding, and clause realization. Details are in
[`docs/generation-api.md`](./docs/generation-api.md). Exact rebuild is not
guaranteed translation accuracy.

Version 0.3.56 adds `composeSequenceSemantics` for full-sequence
`semantic-source/v2` composition with explicit obligations. Details are in
[`docs/sequence-semantic-composition.md`](./docs/sequence-semantic-composition.md).
`semanticFrameFromSequence` remains the first-fragment lift.

Version 0.3.55 adds the EN/DA → Kalaallisut translation contracts
(`validateSourceAnalysis`, `validateMeaningGraph`, `validateTargetPlan`,
`validateGenerationCandidate`) and `adaptSemanticFrameToGrammarianV1`.
Existing `semanticFrameFromSequence` output is unchanged. Shapes, status
axes, and the closed semantic-frame boundary are specified in
[`docs/translation-contract.md`](./docs/translation-contract.md).

Version 0.3.53 adds opt-in structured semantic frames to `presentAnalysis()`
and `analyzeSentence()` when supplied the grammarian ID-first catalog. Legacy
headline/gloss fields remain available; missing source facts yield a nullable
frame rather than a guessed interpretation.

Version 0.3.31 adds soft pedagogy/display stem-ending strength: `resolveStemEndingStrength`, plus a `stemEndingStrength` field (`"strong"`|`"medium"`|`"weak"`|`null`) on preset/seq items from `morphemeEntryToPreset` / `toBuilderItem`. Catalog `morphology_class.strength` stays `strong`|`weak`|`regular` (unchanged); display maps `regular` → learner `medium`. Missing/unknown strength → `null` — never invent from `stem_type`.

Version 0.0.1 cut PWA leakage from the package (oq-api#14):
`buildWord` / `analyzeWord` / conjugation labels take optional `EngineOptions`
(`linguistTerms`, `pronounPreference`, `storage`, `locale`) instead of reading
`localStorage` or oq `kc_exp_*` keys. Defaults need no browser and preserve
the previous stored-default surface forms (linguist terms OFF, pronoun
preference `"he"`). `t()` looks up morphology catalogs only (word-class terms,
conjugation labels, gloss/reason fragments); GUI catalogs stay in oq.

The same 0.0.1 surface also includes the Phase 3 engine extracts (oq-api#15):
custom-morpheme validation, identity-preserving sequence helpers, deconstruct
search extras, dictionary homograph merge, semantic-class/domain catalogs, and
katersat stem/gloss hints. Package version stayed **0.0.1**; oq view wiring waits
for oq#888.

Version 0.0.2 adds generic support for grammarian
`application_logic.special_ending_realizations`, preserving the affix and
following ending as separate analysis-chain morphemes while composing their
published complete surface realization.

Version 0.0.3 fixes the confidence and ranking of those mapped analyses: a
verified complete realization is exact rather than approximate, so it ranks
ahead of unrelated approximate fallback analyses.

Version 0.0.4 assigns mapped affix+ending realizations to a single explicit
surface span in both structured surface results and padded breakdown rows,
while preserving the following ending as a separate zero-width analysis item
and keeping spans ordered and non-overlapping.

Version 0.0.5 corrects per-morpheme ownership within those mapped spans by
using the following ending's citation-form suffix as the generic ownership
anchor.

Version 0.0.6 exports conjugation math (`conjugate`, stem derivation, ending
catalog, `matchConjugationEnding`) that previously lived only behind oq
shims, and drops the dead KalaalliCut / C-port pin from package docs.

Version 0.0.7 adds structured word-splitting markers (`splitWord`) and
`syllabify`. These are syllable / line-break splits, not morpheme spans.

Version 0.0.8 exports the headline picker Deconstruct uses (`headlineGloss`),
`buildMeaningTrace`, Builder next-step grammar (`canFollow` /
`pickMatchingSense` / `INITIAL_STATE`), and dictionary `getAllEntries` /
`suggestWordCompletions`.

Version 0.0.9 exports `injectDictionaryRoots`: turn dictionary headwords
into synthetic `dict_…` stem presets. Off by default — `analyzeWord` is
unchanged unless the caller concatenates the result.

Version 0.0.10 exports `presentAnalysis`: Deconstruct's answer policy
(sequence conjugation, citation reverse-lookup that drops competing cards,
headline, hide bare `dict_*` matches). Opt-in; `analyzeWord` golds unchanged.

Version 0.0.11 adds absolute `url` (frozen API snapshot) and `tarball_url`
(frozen install tarball) to `/api/latest.json`. The JavaScript export list
is unchanged.

Version 0.0.12 exports the remaining helpers a second frontend needs to
consume published data without reaching into `src/*.js`: `flattenGrammarianData`
(the published grammarian URL is `by_id`, while `buildEndingCatalog` wants a
flat array), `loadFullSource` (prime `getAllEntries` / `injectDictionaryRoots`
without a dummy search), plus display labels `formatWordClass`,
`wordClassInfo`, `labelForCategory`, `markedForm`, and `groupCategoriesByType`.

Version 0.0.13 exports the rest of the Builder/Dictionary algorithms a
greenfield client would otherwise copy: search syntax (`compileSearchPattern`,
`matchesSearch`), stem/word-class helpers (`stemWordClass`, `wordClassPath`,
`normalizeWordClassLang`, `WORD_CLASS_SHEET`), Builder category grouping
(`categoryToGroupType`, `capForRender`, `joinRuleForItem`), allomorphy
predicates the Builder paints (`nasalizesPrecedingStop`,
`nasalizesPrecedingConsonant`, `copiesPrecedingVowel`,
`assimilatesAlveolarSchwa`, `anchorSurfaceVariants`), conjugation subject-split
hint (`moodSubjectSplitKey`), and dict priming (`ensureFullLoaded`,
`randomEntry`). `debounce`, data-mirror URLs, IPA/Sumut indexes, related-words,
and color `init()`/`getTree()` stay out — those are PWA chrome or stateful.

Version 0.0.14 publishes human-readable and machine-readable What's New files
for downstream consumers. The JavaScript export list is unchanged. The
discovery payloads advertise the files through a `whats_new` object.

Version 0.3.54 publishes stable consumer feature IDs at `/api/features.json`
and advertises that registry through `/api/latest.json`'s `features_url`.

Stable machine-readable consumer feature IDs and their first available API
versions live in [`.github/api-features.json`](./.github/api-features.json).
Known consumers declare IDs only after using them and link passing integration
tests as evidence. Dated removals belong there too: each needs a replacement
feature and an explicit removal gate before the deadline can authorize removal.

Version 0.1.0 exports the remaining pure output-normalization policies from
oq's Deconstruct view: interactive conjugation selections can be converted
back into structured analysis matches, and analysis sequences can be reduced
to portable Builder-handoff entries.

Version 0.1.7 realizes authoritative infinitive-shaped composition metadata
when structured person endings consume it, preventing glosses such as
`I to have a dog` / `Jeg at have en hund` while preserving person agreement.

Version 0.1.6 fixes context-specific sandhi confidence for cited denominal
verb forms and preserves reverse-search roots when a longer literal root is
morphologically incompatible with the target's ending.

Version 0.2.0 exports a sentence lattice and a checkable clause assembler
(`tokenizeSentence`, `analyzeSentence`, `assembleClause`). Running text is
split on whitespace (enclitics stay on the host). Edge presentation
(Unicode punctuation and emoji) is stripped from each token's analysis
`surface` and kept on `raw` for display. Each token keeps closed
gold/hard_exact readings with band, headline, and mood/case/person/possession;
SOFT_EXACT sits in `also` when a hard reading exists, and is promoted into
`readings` (still banded `soft_exact`) when it would otherwise be a hard zero.
Approximate and open paths are dropped. On `num:{digits}` hosts, plain case
outranks POSS when certainty ties. Pass-through `ENCL_lu` no longer replaces
the verb/NP headline — content is kept and a short “and”/“og” is appended. The
assembler emits one ugly English/Danish clause from a known sentence shape
(intransitive ABS subject bind, transitive ABS object / ERG subject bind, name-is idiom,
subordinate head). Citation-form nouns with no case count as ABS, so they
are not dropped. Exactly two adjacent *citation-form* nouns (null case, no possessor) are
one NP (Danish Fuge-e / English space compound) and then bind:
`sava suppa nerivara` → `I eat sheep soup` / `Jeg spiser fåresuppe`.
A run of three or more citation nouns stays serial
(`sava qimmeq suppa nerivara` → `sheep, dog, soup, I eat him/her/it`).
Morphologically marked ABS does not fuse with its neighbour —
two `N_ABS_SG` plus a verb stay serial. One ABS plus a transitive 3rd-person
object binds the NP as object (`suppa nerivara` → `I eat soup`), never
`I eat him | S/P: soup`, and only when **number matches** (3sg ending +
plural NP, or a 3sg dummy that says `them`, is serial). Nouns split across the verb stay serial. When no
structure rule matches, it fails open: the
literal headlines in token order (`sheep, soup`), never `(no_finite_verb)`
when tokens exist. Empty input is `mode: "empty"`. Danish citation-form
noun sequences compound to one word via a stem-shape Fuge-e heuristic
(`sava suppa` → `fåresuppe`): a monosyllabic
consonant-final modifier that is not s-final takes linking `-e-`. Verbless
Danish compounds any bare ABS (null or marked absolutive, no possessor);
fold-then-bind with a verb is citation-form only (null case). Fuge-e does
not run on a reading that fell through to English (`glossLang: "en"`;
`assembleClause(enLattice, { lang: "da" })` → `sheep, soup`, not
`sheepesoup`). The fallback is tagged on the reading, not sniffed from
the headline string. This is
not the Danish lexicon (`bil`+`hus` → `bilehus`, not `bilhus`).
A citation-form reading sitting in `readings` / `also` wins over a
pedagogical ABS filling for fold and bind; `one soup (the basic form: …)`
strips to `soup`. Serial NP headlines get the same strip (`my knife (one
thing)` → `my knife`). On a transitive verb, marked ABS (case ending,
including possessed) outranks citation: exactly one highest-rank NP binds
as P; leftover citation is `other:` (`sava saviga nerivara` →
`I eat my knife | other: sheep`). Two marked ABS still serial. Two
citation nouns still fold, then bind. `assembleClause` returns
`{ ok, reason, mode, text, parts }` with `mode` one of
`clause | compound | serial | empty`. Serial with tokens is `ok: true` /
`mode: "serial"` — the glosses in `text`, not a closed failure. Empty input
is the only closed failure (`ok: false`, `mode: "empty"`). Not a
translator. Does not invent a copula. Transitive ABS is bound as object when
the ending’s object is 3rd person and number matches; otherwise serial. #138 steps 6–7.

Version 0.3.6 documents consumer-visible assembler and dictionary matching
already used by Misiliineq: citation `!`/`?` on `tuavi!` is display-only
(`citationPunct`); sentence runs exact-match Oqaasileriffik+Katersat tokens
without the full-dictionary checkbox; marked ABS outranks citation on a
transitive verb.

Version 0.3.50 keeps bare numeric dates digits-only in serial text (no invented
`den`/`on`, no month-name expansion) and drops a leading `og`/`and` from an
ENCL_lu Abs when the filled open-slot gloss already opens with `fordi`/`because`.

Version 0.3.51 joins serial `assembleClause` parts from source punctuation:
sentence-final `.`/`!`/`?` (and hard newlines between tokens) cut sentences;
`,` cuts clauses; otherwise `", "` within a clause. Middot (` · `) is not used
for serial chip joins (sense merges in morpheme-meta still use middot). Misiliineq
clause-part chips prefer each part's `joinAfter` so the strip matches `.text`.
If a part's text already ends with `,` / `.` / `!` / `?`, `joinAfter` is normalized
so a second clause comma (or stacked sentence mark) is not emitted.

Version 0.3.11 generalizes gloss composition for subject-bearing display shorts: stemOut prefers subject-less composition/verb_phrase companions, English ___-ing/___-s/past skips closed-class heads, ABS drops stacked inner “the”, and headlineGloss prefers plain shortGloss over pedagogy-parenthetical dumps (api.oq.gl#121).

Version 0.3.10 accepts published proper-name hosts with a single leading capital as exact reverse-search surfaces (Alixa), while still rejecting all-caps/interior uppercase catalog noise (api.oq.gl#214).

Version 0.3.9 ships post-v0.3.8 consumer fixes already on master: longer lexical-span ranking (#199), dict citation seal + declined prefix injection (#200), soft-demote intensifier/future junk stacks (#201), sentence edge punctuation/emoji strip (#210), misiliineq unparsed-token styling (#211), and stress follow-ups — presentation raw restore, digit LOC/INS preference, soft-also promote, ENCL_lu headline (#212).

Version 0.3.8 adds Arabic-digit case hosts (`num:{digits}` + case, including m/n fold for `21ni` / `5nik`; api.oq.gl#197). Version 0.3.7 documents JOIN_ASSIMILATIVE uvular halfs already used by the
morph engine: before m-initial continuers `q+m→rm` (velar `k+m→mm` kept),
and before l-initial continuers / `ENCL_lu` `q+l→rl` (non-uvular `C+l→ll`
kept). Shared `assimilativeReplaceFinal` in `applyStemOp` and stage-3.


## Pages distribution

[https://api.oq.gl](https://api.oq.gl) (Cloudflare Pages) is the canonical
host. [https://jandahl.github.io/oq-api/](https://jandahl.github.io/oq-api/)
publishes the same `site/` archive and is not a separate compatibility line.
Frozen patches stay at `/api/vX.Y.Z/` and `/downloads/oq-api-X.Y.Z.tgz`. Each
`X.Y` line also gets a rolling copy. The current package line is 0.4:

```
https://api.oq.gl/api/v0.4-latest/public-api.js
https://api.oq.gl/downloads/oq-api-0.4-latest.tgz
```

`v0.4-latest` tracks the newest 0.4.x and is overwritten every deploy. Older
lines keep their own alias (`v0.0-latest` through `v0.3-latest`) and
do not move when 0.4 ships a patch. Pin `/api/v0.4.Z/` or
`oq-api-0.4.Z.tgz` (the frozen URL in `latest.json`, not a patch copied from
an old document) when you need a reproducible build. Do not pin
`0.4-latest.tgz`. The 0.3 line stays frozen at its last published snapshot.

Pages replaces the entire site on each publish. Frozen snapshots and tarballs
are rebuilt from **git tags** (`vX.Y.Z`). After a successful deploy, Pages
tags HEAD as `v<API_VERSION>` if that tag is missing, so the next publish
still has a pin to rebuild. Do not pin a rolling `X.Y-latest` tarball.

Discovery (additive fields on the existing `latest.json` payload; new
`versions.json` index):

| URL | Payload |
|---|---|
| [`/api/latest.json`](https://api.oq.gl/api/latest.json) | Current HEAD: `version`, `tag`, `url` (frozen snapshot), `tarball_url` (frozen tarball), `api_url`, `release_url`, `whats_new`, `features_url`, plus `line` / `line_url` / `module` for the rolling alias. |
| [`/api/features.json`](https://api.oq.gl/api/features.json) | Stable API feature IDs, their first available versions, known published consumers, and any dated deprecation/removal gates. |
| [`/api/versions.json`](https://api.oq.gl/api/versions.json) | Every frozen patch and every `vX.Y-latest` line alias, each with `whats_new` links. |
| [`/whatsnew/0.4-information.md`](https://api.oq.gl/whatsnew/0.4-information.md) | Human-readable changes relevant to consumers on the current 0.4.x line. |
| [`/whatsnew/0.4-information.json`](https://api.oq.gl/whatsnew/0.4-information.json) | The same entries as structured data: `{ schema_version, api_line, entries }`. |

The files are cumulative per `X.Y` compatibility line. Each entry identifies
the exact `X.Y.Z` release that introduced the change, and newest releases sort
first. Older lines use the same `/whatsnew/X.Y-information.{md,json}` pattern.

The live example at [`/examples/live-import.html`](https://api.oq.gl/examples/live-import.html)
imports `v0.4-latest` and lists these URLs at the bottom.

## Stability posture

**There is no stability promise yet — deliberately.** The morphology engine is
under heavy active development: the sandhi taxonomy is only partially
implemented, `OUT_OF_SCOPE_IDS` churns weekly, and the upstream grammarian
schema still moves. While `API_VERSION` is `0.x`, any commit may rename,
reshape, or drop any export. What v0 *does* give you:

- **One blessed entry point.** Import from `src/public-api.js` only, or from
  the package subpaths `oq-api/engine`, `oq-api/catalog`, and
  `oq-api/generation`. Subpath indexes re-export the direct modules and do
  not clone; the package root still does. Everything else under
  `src/engine/`, `src/catalog/`, and `src/generation/` — including the other
  exports in `morph-engine.js`, `allomorphy.js`, `phonology.js`,
  `morpheme-meta.js`, `dict-data.js` — is internal plumbing, free to change
  without notice. Deep imports of a published snapshot
  (`/api/v0.4-latest/engine/morph-engine.js`) follow the folder layout and
  break when that layout moves. `v0.3-latest` stays the previous frozen line.
- **Deliberate drift.** `test/public-api.test.js` snapshots the export list;
  changing the surface requires updating the module, this document, and the
  snapshot together, plus an `API_VERSION` bump.
- **Detectability.** `API_VERSION` tells a consumer which surface it got.

A real 1.0 promise is gated on the stabilization criteria discussed in
oq#394 (sandhi taxonomy substantially complete, engine-behavior churn
flattened). This module is **not published to npm**; consume it by importing
the file directly. It is framework-free and runs both as a browser ES module
and under Node (≥ the version CI uses; no bundler, no build step).

## Quick start

```js
import {
  buildWord, morphemeEntryToPreset, findExactDictMatch,
} from "./public-api.js";

// entries: flattenGrammarianData(await fetch(GRAMMAR_MORPHEMES_URL).then(r => r.json()))
const seq = [stemEntry, affixEntry, endingEntry]
  .map((e) => morphemeEntryToPreset(e).seq[0]);

const { ok, word, approximate, closed } = buildWord(seq);
// → { ok: true, word: "qimmeqarpoq", approximate: false, closed: true, ... }

const dictHit = await findExactDictMatch(word); // entry object or null
```

## Surface

### Versioning

| Export | Contract |
|---|---|
| `API_VERSION` | Semver-shaped string, `0.x` while experimental. MINOR bump on any surface change, PATCH on doc/behavior-only fixes. |

### Word building — composer

| Export | Contract |
|---|---|
| `buildWord(seq, options?)` | Full pipeline in one call; returns `{ ok, word, approximate, closed, error, errorKey, errorAt, reason, details }`. `errorKey` and `details` are stable machine-readable diagnostics; do not parse `reason`. `options` is the same optional `EngineOptions` bag as conjugation labels; build itself does not read settings or storage. |

### Word building — engine-level pieces (advanced)

For consumers that need incremental building or their own pipeline wiring.

| Export | Contract |
|---|---|
| `WordBuilder` | Incremental engine: `init()`, `add(morpheme)` → error code, `removeLast()`, `getWord()` (raw, un-respelled), `getError()`, `hasSandhiConflict()`. |
| `ROOT`, `AFFIX`, `INFLECTION`, `ENCLITIC`, `DERIVATIONAL_ENCLITIC` | Sequence-item `type` values for `WordBuilder.add()`. |
| `JOIN_NONE`, `JOIN_TRUNCATIVE`, `JOIN_ASSIMILATIVE`, `JOIN_ADDITIVE` | Coarse join operations. Rule values coming out of `morphemeEntryToPreset` may also be engine-internal sentinels — treat any rule value you didn't construct yourself as opaque. |
| `MORPH_OK`, `MORPH_ERR_*` (7 codes) | Error codes from `WordBuilder.add()`/`getError()` and `buildWord().error`. |
| `MAX_WORD_LENGTH`, `MAX_MORPHEMES`, `MAX_MORPHEME_LENGTH` | Engine limits. `MAX_MORPHEMES` (24) is the hard sequence ceiling for `buildWord` / Deconstruct; default reverse-search depth stays `DEFAULT_MAX_MORPHEMES` (8). |
| `applyAllomorphy(seq)` | Tier-2 pre-pass: returns a new array with each item's `.text` rewritten to its correct surface allomorph given its neighbours. Pure. |
| `hasApproximateMorpheme(seq)` | Whether any item needs a process the engine doesn't implement (should be flagged approximate). |
| `nasalizesPrecedingStop(morpheme)` | True when this morpheme nasalizes a preceding stop (`ENCL_aa` family). Builder paints this as a join hint. |
| `nasalizesPrecedingConsonant(morpheme)` | True for `ENCL_guuq`. |
| `copiesPrecedingVowel(morpheme)` | True for `N_uneq_N`. |
| `assimilatesAlveolarSchwa(morpheme)` | True when this morpheme assimilates a host-final alveolar schwa (`N_cuaq_N`). |
| `anchorSurfaceVariants(morpheme)` | Citation/allomorph spellings a Builder search should accept as this morpheme, or `null`. |
| `validateSequence(seq)` | Word-class (N/V) grammar check: `{ valid, state, errorAt, reason }`. |
| `INITIAL_STATE` | Empty morphotactic state (`category: null`, `closed: false`). Start of a Builder sequence. |
| `canFollow(state, morpheme)` | Whether `morpheme` may attach after `state`. `{ ok, reason?, reasonKey? }`. Unannotated free-form items pass. |
| `pickMatchingSense(state, morpheme)` | Homograph sense whose `category_shift.from` matches `state.category`, or `null`. |
| `respellSurface(word)` | Whole-word display respelling (q→r, vowel openness, ŋ→ng). Apply once, after building — never inside the join pipeline. |
| `TYPE_INT` | `{ ROOT: 0, AFFIX: 1, INFLECTION: 2 }` — numeric `type` values for `buildCustomMorphemeItem`. Same numbers as `ROOT` / `AFFIX` / `INFLECTION`. |
| `buildCustomMorphemeItem(values, options?)` | Pure custom-morpheme validation. `{ ok: true, item }` or `{ ok: false, reason }`. No DOM. `options` is the uniform `EngineOptions` bag (unused by the builder itself). |
| `moveSeqItem(seq, from, to)` | Identity-preserving reorder: returns a new array with the item at `from` moved to `to`. Non-arrays become `[]`; out-of-range / non-integer indices return a shallow copy. |
| `isBuilderItem(item)` | True when `item` is already `toBuilderItem()`-shaped (the identity `buildWord` preserves). |
| `compileSearchPattern(query, advanced?)` | Limited search syntax → `RegExp` or `null`. `advanced: true` treats `query` as a regex. Ordinary punctuation is literal. |
| `matchesSearch(text, pattern)` | Case-insensitive test; resets `pattern.lastIndex`. False when `pattern` is null. |
| `joinRuleForItem(item)` | `{ rule, fromLeftSandhi }` — the join `WordBuilder.add` will actually run. Prefer this over reading `item.join` when `left_sandhi` is mapped. |
| `resolveStemEndingStrength(input)` | Soft pedagogy/display stem-ending strength: `"strong"` \| `"medium"` \| `"weak"` \| `null`. Accepts a seq/preset item (reads `.morphology_class`) or a bare `morphology_class`. Catalog `strong`/`weak` pass through; catalog `regular` → learner `medium` (display only — not a third phonological class). Missing or unknown strength → `null`; never invents from `stem_type`. Raw `morphology_class.strength` stays intact. |
| `categoryToGroupType(groups)` | Map category key → group type from `groupCategoriesByType()` output. |
| `capForRender(items, limit?)` | `{ shown, overflow }` slice for Builder lists. Default limit 300. |

### Word-class colors

| Export | Contract |
|---|---|
| `WORD_CLASS_THEMES` | Serializable canonical `default` and `light` theme choices. |
| `getWordClassColors(classPath, theme?)` | Purely returns `{ border, fill, text }` for a hierarchy path using the supplied theme, without global state or initialization. |
| `formatWordClass(desc, opts?)` | Human label for a raw dictionary word-class string (`"taggit"`, `"v"`, `"oqaluut susalik"`). `opts: { lang?: "en"\|"da"\|"kal"\|"both", abbrev?: boolean }`. Unknown labels pass through. |
| `wordClassInfo(desc)` | Registry row for that label (`{ id, en, da, kal, abbrEn, classPath, … }`), or `null`. |
| `stemWordClass(desc)` | `"N"` / `"V"` / `null` for whether this label can seed a Builder root. Blank → `"N"`. |
| `wordClassPath(desc)` | Hierarchy path array for that label, or `[]`. |
| `normalizeWordClassLang(value)` | `"kal"` / `"en"` / `"da"` / `"both"` (default for anything else). |
| `WORD_CLASS_SHEET` | Static cheat-sheet rows (`{ id, en, kal, da, classPath, … }[]`). |

### Deconstruct — word analysis (the inverse of `buildWord`)

Analysis matches also expose `reliability` (`attested`, `compositional`, `approximate`, or `surface_only`), `qualityClass` (`exact_attested`, `exact_compositional`, `approximate`, or `surface_only`), `presentationSafe`, and `qualityReasons`; excluded candidates have `qualityClass: "rejected"`. The result's `quality` object reports the top ranked candidate summaries, excluded unsafe chains and their reasons, and whether configured search limits may have hidden candidates. `certainty` remains the existing numeric 0–65535 confidence score.

`glossSummaryItems()` includes a `semanticStep` diagnostic per item: `{ kind, input, output, safe, languageSource }`. It describes the mechanical gloss-threading operation (root gloss, identity, composed template, fallback, or unresolved step) and reports whether the row fell back to English, Danish, or a scholarly `meaning` field. It is not an attestation or a guarantee that the output is idiomatic. Consumers should keep per-morpheme glosses distinct from a whole-word translation and must not present unsafe steps as certain. `headlineGloss(items, { requireLanguage: true })` withholds a headline that would silently carry an English gloss into Danish (or vice versa); the default preserves legacy fallback behavior but returns `languageSource` for explicit labeling.

| Export | Contract |
|---|---|
| `analyzeWord(word, presets, opts?)` | Synchronous search returning verified matches, sorted by `certainty` (desc). Each match includes `certainty: number` (0…65535), `confidence: "exact" \| "approximate"`, `band: "gold" \| "hard_exact" \| "soft_exact" \| "approximate" \| "floor"`, `rankReasons: string[]`, and `certaintyFactors: { key, delta }[]`. `confidence` stays `"exact"` for SOFT_EXACT (back-compat); translator-facing consumers should read `band`. The result also includes `budget: { maxMorphemes, maxNodes, beamWidth, hitMorphemeCap, hitNodeCap, hitBeamCap, stopReason }` so a miss can distinguish a depth cap from a catalog gap or a node-budget exhaustion — default depth stays 8; callers raise `opts.maxMorphemes` (up to `MAX_MORPHEMES` = 24) rather than the engine silently doing so. Optional `opts.workedExampleChainKeys` (Set/array of `"id id …"` chains) snaps gold; optional `opts.katersatLexicalClasses` (Map/Record of stem id/text → Katersat class codes) supplies soft Katersat knobs — omit for deterministic `n/a`. Published grammarian `semantic_tags` / `enclitic_hosts` gate nonproductive enclitics on the hard path and feed `productivity_host_ok` inside HARD_EXACT. A lexical stem that shares an ending with a longer composition (2+ extra prefix morphemes) demotes that composition to SOFT_EXACT. Surfaces matching `^\d+` + case remainder inject a synthetic `num:{digits}` noun host (api.oq.gl#197), with oblique m/n fold (`mi`↔`ni`, …). |
| `analyzeWordAsync(word, presets, opts?, { signal }?)` | Same search and match shape as `analyzeWord`, yielding to the event loop between chunks so it never blocks rendering; `signal` (an `AbortController`'s `.signal`) cancels a stale search, rejecting with a `DOMException` named `"AbortError"`. |
| `cacheAnalysisResult(word, result)` | Seeds the exact-match cache with a precomputed result (e.g. from completion generation), so a later `analyzeWord`/`analyzeWordAsync` call for the same word returns it directly regardless of `opts`. |
| `computeMorphemeBreakdownRows(items, word, seq)` | Pure, DOM-free layout computation behind oq's own per-morpheme breakdown table: given `glossSummaryItems(seq)`'s `items`, the built `word`, and the (unfiltered, including any `Ø` items) `seq` that built it, returns one row per non-`Ø`, glossed item with `{ item, label, marker, j, text, changedRanges, leftPad, rightPad, surfaceLeftPad, surfaceEnd, surfaceText, surfaceRightPad }`. `text` is the morpheme's own plain declared/citation spelling (never allomorph- or sandhi-resolved — e.g. `"vunga"`, not `"punga"` or `"rpunga"`); `changedRanges` (`{start, end}[]`, indices into `text`) marks which of its own letters don't survive unchanged into the real word — both a LEADING change (its own allomorph pick differing from citation, e.g. `"-vunga"` picking `"-punga"`) and a TRAILING one (the NEXT morpheme's boundary altering its tail, e.g. `"-qaq"`'s own final `"q"` becoming `"r"`) are covered uniformly, including identity swaps (a letter replaced by a different one), not just outright deletions. `leftPad`/`rightPad` are dot-padding counts around `marker+text` that sum to exactly `word.length`, right-anchored against the next row's real start (or the word's end, for the last row) whenever `text` is shorter than the real width its own boundary added; `surfaceLeftPad`/`surfaceEnd`/`surfaceText` describe the row's real, fully-resolved span within `word` directly. Column positions are derived from replaying the real build pipeline (not a naive prefix diff), correctly handling cases a naive diff gets wrong (e.g. a join that inserts a surface letter belonging to neither morpheme's own declared text). |
| `resolveMorphemeSurfaces(seq, word)` | Plain structured sound-change data with no UI/padding concerns and no `items`/gloss dependency — just the `seq` that built `word` and the `word` itself. Returns one record per real, non-`Ø` entry (unlike `computeMorphemeBreakdownRows`, this includes a glossless entry too — a display choice that function makes, not a structural fact this one inherits) with `{ j, id, marker, citationText, resolvedText, changedRanges, surfaceStart, surfaceEnd, surfaceText }`. `surfaceStart`/`surfaceEnd`/`surfaceText` are the morpheme's real, final resolved spelling and its exact span within `word` — the same ground truth oq's own "sound change" column shows, as plain data: every record's `surfaceText`, concatenated in order, reconstructs `word` exactly, with no gap or overlap. `changedRanges` is the same generalized sound-change marking `computeMorphemeBreakdownRows` exposes (see above), located within `citationText`. |
| `splitWord(word)` | Oqaasileriffik syllable / line-break splits as **data**, not an invisible display trick. Returns `{ word, syllables, breaks, hyphenated, visible }`. `word` is the input unchanged. `syllables` are the parts; `breaks` are offsets in `word` after which a break is legal; `hyphenated` inserts U+00AD for CSS wrap; `visible` joins with `-` for teaching / copy / logs. Orthogonal to `resolveMorphemeSurfaces`. Never mutates `buildWord().word`. Independent reimplementation of the described algorithm; no license is asserted over Oqaasileriffik's page. |
| `syllabify(word)` | Thin alias for `splitWord(word).hyphenated` — the form oq's Builder currently paints. Prefer `splitWord` in new consumers. |
| `presentAnalysis(query, analyzeResult, opts?)` | Opt-in Deconstruct policy. `{ matches, conjugation, headline, completions }`. Sequence conjugation (exact card, last item in `opts.catalog`) wins; else citation reverse-lookup (`lookupCitation("${prefix}vaa")`, then a capped scan of `opts.citations`, never a hidden full-dict loop) and **drops** competing `matches`. `hideBareDictRoots` default on. `analyzeWord` itself is unchanged. Pass `catalog` (from `buildEndingCatalog`) and optional `stemHints`. Passing `semanticFrameCatalog` (raw grammarian ID-first catalog) adds nullable `semantic_frame` values to returned matches while preserving existing headline and fallback behavior. |
| `builderSequenceEntries(seq)` | Pure normalization for handing an analysis sequence to a Builder. Catalog-backed non-empty items become ids; anonymous items retain their text, numeric type/join, and morphology fields. Zero-surface (`""` or `"Ø"`) endings are omitted because they are useful analysis evidence but produce unusable blank Builder rows. Non-arrays return `[]`; inputs are not mutated. |
| `synchronizeConjugationMatch(presentation, match, result)` | Purely applies an interactive conjugation result to an analysis match. Returns a new match with `word`, `approximate`, and a catalog-backed structured ending after `presentation.prefixSeq`, preserving useful metadata from `presentation.baseMatch` (or `match`). An incomplete result returns `match` unchanged. This is the same output policy oq's Deconstruct conjugation foldout uses. |
| `DEFAULT_MAX_MORPHEMES` | Default `opts.maxMorphemes` for `analyzeWord` / `analyzeWordAsync` (8). Callers raise up to `MAX_MORPHEMES` (24) for long words; the default is not raised globally. |
| `DEFAULT_NODE_BUDGET` | Default `opts.maxNodes` for `analyzeWord` / `analyzeWordAsync` (100000). oq's settings chrome may override it; this package does not read that setting. |
| `isBareDictRootMatch(match)` | True when a deconstruct match is a dictionary-injected root plus empty continuers. |
| `suggestFuzzyRoots(word, presets, opts?)` | Edit-distance root suggestions when analysis finds no exact parse. `opts`: `{ maxDistance?, maxSuggestions? }` plus `EngineOptions`. |
| `suggestMorphCompletions(prefix, presets, maxCount?, options?)` | Prefix completions from the morpheme catalog (closed built roots). |


### Morpheme data — grammarian entries → presets

| Export | Contract |
|---|---|
| `morphemeEntryToPreset(entry, opts?)` | Maps one raw grammarian/katersat-schema morpheme entry to a preset; `preset.seq[0]` is the sequence item to feed `buildWord`. Pure. |
| `mergeMorphemeSources(results, sources)` | Merges multiple fetched morpheme sources into one deduplicated preset list: `{ presets, anyOk, failed }`. |
| `toBuilderItem(item)` | Preset `seq[]` item → numeric-typed engine item. Idempotent on already-numeric items. |
| `glossSummary(seq, opts?)` | Human-readable gloss chain for a sequence, as an array of pre-joined `"${marker}${text} — ${gloss}"` strings (one per non-filtered item — see `glossSummaryItems` below for what "filtered" means and why nothing is filtered here). `opts: {lang, showOther}` (oq#409, both optional, defaulting to `en`/`never` — the pre-#409 behavior) select the display language for each item's `plainGloss` and whether the non-preferred language is surfaced when it differs. Joins on each item's scholarly `gloss` field — for the shorter, single-sense phrasing oq's own UI actually renders (`shortGloss`/`rawShortGloss`), call `glossSummaryItems` directly instead (oq#824). |
| `glossSummaryItems(seq, opts?)` | The structured form `glossSummary` reduces to strings: one item per input morpheme, `{ marker, text, gloss, shortGloss, rawShortGloss, meaning, preset, stemIn, stemOut, moodLabel, valencyIn, valencyOut, valencyEffect, htrRole }`. `gloss` is the raw, scholarly, potentially multi-sense text; `shortGloss` is the single-sense, blank-filled phrasing oq's own Deconstruct/Word Builder UI renders; `rawShortGloss` is the same but with the composed stem left as a literal `"___"` instead of filled in — pick whichever fits your UI. `marker` is `""` for a root, `"+"`/`"-"` for an additive/truncative-joined continuer, or the literal string `"Ø"` for a null/zero ending (`morpheme-meta.js`'s own `marker = isRoot ? "" : !text ? "Ø" : ...`) — a `"Ø"` item's `text` is always empty, since a zero ending carries no real bound-morpheme spelling of its own. **Not filtered out for you**: oq's own rendered breakdown (`docs/morpheme-breakdown.js`'s `renderMorphemeBreakdown`) drops `marker === "Ø"` items before displaying a sequence, and most consumers building a similar per-morpheme breakdown display will want to do the same — but this function returns every item honestly, since a caller auditing the full sequence (not just displaying it) may need the Ø entries too. Same `opts` as `glossSummary`. |
| `resolveGlossText(plainGloss, meaning, lang, showOther)` | Resolves one morpheme's raw `plain_gloss` (`{en, en_short, da, da_short}`, any key optional) plus its scholarly `meaning` fallback into `{ text, tag, secondary }` for the given language preference and show-other mode (oq#409). |
| `headlineGloss(items, opts?)` | Composed headline for a `glossSummaryItems()` list — not always the last item. Walks backward past unfilled `"___"` and past Ø identity endings whose gloss only repeats `stemIn`. Prefers a structured `en_short`/`da_short` object. Strips a mood label already shown as a pill. `opts: { lang?, locale?, moodLabel? }`. Identity skip is locale-specific (`en`: leading `"one "`, `da`: leading `"et "`/`"en "`); unknown locales skip that filter. Returns `{ text, item, moodLabel }`. |
| `stripRedundantMoodLabel(text, moodLabel)` | Strip a leading `"label: "` / `"label "` when the mood pill already carries that label. |
| `buildMeaningTrace(text, items, last?)` | Colour-span ranges `{ start, end, seqIndex }[]` locating each item's `meaningContribution` in the composed sentence. Paint stays with the caller. |
| `flattenGrammarianData(data)` | Recover a flat entry array from the published grammarian `by_id` export used by `GRAMMAR_MORPHEMES_URL`. Empty / unknown payloads return `[]`. Feed the result to `morphemeEntryToPreset` or `buildEndingCatalog`. `mergeMorphemeSources` already does this internally. |
| `labelForCategory(cat, options?)` | Locale-aware label for a raw category key (`"noun_cases_possessive"`). Falls back to de-underscored text when the morphology catalog has no entry. `options.locale` selects the catalog. |
| `markedForm(preset)` | Citation with Greenlandic boundary marker: stem unmarked, additive continuer `+form`, truncative/assimilative `-form`, null ending `Ø`. |
| `groupCategoriesByType(morphemeData)` | Bucket distinct `category` values under morpheme-type groups (`stem`, `derivational_affix`, `inflectional_ending`, …). Empty groups omitted. |
| `SCHEMA_MAJOR_VERSION` | The grammarian data-schema MAJOR version retained for compatibility with existing consumers. |
| `BY_ID_SCHEMA_MAJOR_VERSION` | The grammarian by-ID export schema MAJOR version consumed by this surface. |
| `GRAMMAR_MORPHEMES_URL` | Version-pinned URL of the published grammarian `morphemes-by-id.json`. |
| `checkSchemaVersion(meta, expectedMajor?)` | Returns a warning string on a real MAJOR mismatch in fetched data's `meta`, else `null`. |

### Semantic frames — Grammarian Issue #648

| Export | Contract |
|---|---|
| `SEMANTIC_FRAME_SCHEMA` | The canonical schema identifier, `https://github.com/jandahl/oq-grammarian/semantic-frame/v1`. |
| `loadMorphemesById(options?)` | Fetches the authoritative ID-first `by_id` catalog. Tries Cloudflare and then the GitHub Pages mirror by default; callers may provide `urls` and `fetchImpl`. |
| `resolveMorphemeReference(reference, catalog)` | Resolves an ID or entry against `catalog.by_id` without decoding the ID. |
| `resolveConstructionReference(reference, catalog)` | Resolves an inline construction object or an opaque catalog reference; IDs are never interpreted. |
| `semanticFrameFromSequence(sequence, catalog)` | Lifts a predicate fragment and reusable inflection construction into the language-neutral v1 IR. It returns `null` when required facts are absent, never infers from glosses, and never includes finite target-language realizations. |
| `semanticFrameOrLegacy(sequence, catalog, legacy?)` | Returns the structured frame when available, otherwise the caller’s existing legacy analysis/gloss. |
| `attachSemanticFrames(result, catalog)` | Adds optional `semantic_frame` values to existing analysis matches without changing legacy fields. |

The IR preserves source provenance and represents `implicit_action` as an
unresolved slot. Context-dependent rendering is intentionally outside this
module. `tools/measure-semantic-frames.mjs` reports catalog size and repeated
lookup/build timing before any consideration of SQLite or incremental indexes.

`SEMANTIC_FRAME_SCHEMA` is the full URI written by this package.
`SEMANTIC_FRAME_GRAMMARIAN_VERSION` is the short `semantic-frame/v1` token
required by the closed Grammarian JSON Schema. They are not interchangeable.
`adaptSemanticFrameToGrammarianV1(frame)` rewrites only that identifier and
returns the closed document, or `{ ok: false, frame: null, errors }` when the
frame has a non-contemporative mood, lacks `arguments.object`, or carries a
property the closed schema does not allow. It does not invent an object,
change mood, or drop extra fields. Callers that cannot accept a refusal keep
the original API frame.

### Public sentence generation — epic #280 / #292

Versioned orchestration over interpretation (#288), grounding (#289), and
clause realization (#291). See [`docs/generation-api.md`](./docs/generation-api.md).

| Export | Contract |
|---|---|
| `GENERATION_RESULT_SCHEMA` | Result id `generation-result/v1`. |
| `GENERATION_API_FEATURE` | Feature id `sentence-generation` for api-features. |
| `generateFromMeaning({ meaningGraph, resources, options? })` | Async generation from a supplied meaning graph. No parser service. Injected `morphCatalog` + `lexicalIndex` (+ optional crosswalks/revisions). Cancellation via `AbortSignal`. |
| `generateFromSource({ text, language, provider, resources, options? })` | Async EN/DA orchestration with an injected source-syntax provider. Provider failure is explicit; dictionary lookup stays labeled lookup; no silent external translation. |
| `generationCacheKey(...)` | Cache identity over source/context and parser/lexicon/grammar/schema/engine/crosswalk revisions. |
| `generationSupportedCoverage(inventory?)` | Honest coverage notes; exact rebuild ≠ translation accuracy. |
| `scoreGenerationCandidate(...)` | Mechanism ranking factors only. Approximate never outranks exact on the morphology axis. |

Hard coverage/validity gates run before ranking. Truncation reporting is
separate from unsupported meaning. Competing/unsupported source readings stay
visible. No whole-input phrase templates.

### Generation sentence cards — epic #280 / #293 / #295

Projects `generation-result/v1` into the shared sentence/deconstruction card
model. Retained sequences are never re-analyzed. See
[`docs/generation-card.md`](./docs/generation-card.md). Soft-must: finished
UI state gates on honesty via `generationCardFinishedState` — never
`card.ok` alone; `projection_kind` keeps recorded-chain projections labeled
distinct from live generation.

| Export | Contract |
|---|---|
| `GENERATION_CARD_SCHEMA` | Card id `generation-card/v1`. |
| `GENERATION_CARD_FEATURE` | Feature id `generation-sentence-cards`. |
| `GENERATION_CARD_FINISHED_SCHEMA` | Finished-state helper id `generation-card-finished/v1`. |
| `GENERATION_CARD_PROJECTION_RECORDED` | Projection kind `recorded_chain`. |
| `GENERATION_CARD_PROJECTION_LIVE` | Projection kind `live_generation`. |
| `GENERATION_CARD_PROJECTION_UNKNOWN` | Projection kind `unknown`. |
| `GENERATION_DISPLAY_CONTROL_KINDS` | Frozen list of presentation-only control kinds. |
| `GENERATION_MEANING_CONTROL_KINDS` | Frozen list of controls that change meaning and require regeneration. |
| `indexMorphCatalog(catalog)` | Normalize array / id-map catalogs for sequence resolution. |
| `resolveSequenceItems(ids, catalog)` | Retained morpheme ids → builder `seq` items; missing ids stay visible. |
| `projectGenerationCard(result, options?)` | Project a generation result (or ranked primary) into card data with rebuild honesty, builder handoff, listening surfaces, `projection_kind`, and expandable debug. Requires `morphCatalog` to resolve retained ids. `ok` means showable (includes partial); use `generationCardFinishedState` for finished success. |
| `projectGenerationWordCard(args)` | One retained chain → AnalysisMatch-compatible card word. |
| `projectRecordedChainCard(args)` | Project a corpus `recorded_chains` engine rebuild into the same card shape (`projection_kind: recorded_chain`; pending_review; not gold / not live generation). |
| `inferGenerationCardProjectionKind(resultOrCard)` | `recorded_chain` vs `live_generation` vs `unknown`. |
| `generationCardFinishedState(card)` | Honesty gate for finished UI: `complete` only when live, exact, full coverage, reviewed evidence, no open obligations. Never treat `card.ok` alone as finished success. |
| `generationCardDebugPayload(result, ranked?)` | Expandable parser/derivation evidence (not the primary claim). |
| `displayPreferencesFromMeaning(graph, lang?)` | Seed pronoun/number/determination display prefs from meaning. |
| `classifyGenerationControlEdit(edit)` | Separate display-only controls from meaning edits that require regeneration. |
| `generationExampleViewsFromCase(case)` | Direction-preserving EN→KL / DA→KL views from one semantic case. |
| `listModalGenerationExampleViews()` | All modal cases as direction-specific views. |
| `generationCardsEquivalentForReview(a, b)` | Same retained plan/card across EN/DA (dog canary). |

### Clause / multiword realization — epic #280 / #291

Internal stage `clause-realization/v1` (`src/generation/clause-realization.js`,
`docs/clause-realization.md`). Extends `realizeWordPlans` (#290) with
incorporation vs independent arguments, instrumental quantity companions,
clause-link inventory checks, and AITWG Def 8.6 word order. Consumed by the
public generation API (#292); not a separate public export by itself.

### Sequence semantic composition — epic #280 / #286

Forward composition over a complete selected morpheme sequence. Consumes
Grammarian `semantic-source/v2`. Does not certify a chain from one recognized
fragment. See [`docs/sequence-semantic-composition.md`](./docs/sequence-semantic-composition.md).

| Export | Contract |
|---|---|
| `SEQUENCE_COMPOSITION_SCHEMA` | Result id `sequence-composition/v1`. |
| `SEMANTIC_SOURCE_CONTRACT` | Source id `semantic-source/v2`. |
| `normalizeSequenceSelection(item, index)` | Normalize an id string or selection object (`occurrence_id`, `sense_id`, `variant_index`, `revision`, `reading`). |
| `resolveCatalogVariant(reference, catalog)` | Resolve `by_id` without arbitrarily selecting the first of several variants. Ambiguous ids stay unresolved until `variant_index` or a unique `revision` is supplied. |
| `composeSequenceSemantics(sequence, catalog, options?)` | Fold lexical senses, typed operators, incorporation, and inflection (including zero-surface endings). Returns obligations for missing semantics, references, or bindings. `options.mode` is `"generation"` (default) or `"analysis"`. `semantic_frame_v1` is always `null`; call `semanticFrameFromSequence` separately for the historical fragment lift. |

### Translation contracts — EN/DA → Kalaallisut

Versioned documents for later parsing and generation. Validators reject
dangling ids, contradictory bindings, and missing required fields. They do
not parse text or build words. Full field lists and the migration decision
are in [`docs/translation-contract.md`](./docs/translation-contract.md).

| Export | Contract |
|---|---|
| `TRANSLATION_CONTRACT_VERSION` | Family id `translation-contract/v1`. |
| `SOURCE_ANALYSIS_SCHEMA` | `source-analysis/v1`. |
| `MEANING_GRAPH_SCHEMA` | `meaning-graph/v1`. |
| `TARGET_PLAN_SCHEMA` | `target-plan/v1`. |
| `GENERATION_CANDIDATE_SCHEMA` | `generation-candidate/v1`. |
| `validateSourceAnalysis(document)` | `{ ok, errors }`. `errors` are `{ code, path, message }`. English or Danish only. Grammatical relations stay here; semantic roles do not. |
| `validateMeaningGraph(document)` | Source meaning. No Kalaallisut surface or morpheme id. Polarity is an explicit operator. An indefinite article is not an exact quantity. Ambiguous person/number must keep its candidate readings. |
| `validateTargetPlan(document, context?)` | Lexical and construction choices. `context.meaningGraph`, when supplied, rejects dangling meaning ids and any entity or event the plan drops. Unresolved choices cannot carry an entry. Repeated morphemes need distinct `occurrence_id`s. |
| `validateGenerationCandidate(document, context?)` | One outcome: `generated`, `missing_resource`, `unsupported`, `ambiguous`, `truncated`, or `cancelled`. Morphology, meaning coverage, evidence/review, and search completion are separate. Approximate or incomplete morphology cannot be `attested_surface`. `context.targetPlan` and `context.knownCandidateIds` check dangling ids. |
| `SEMANTIC_FRAME_GRAMMARIAN_VERSION` | Closed-schema token `semantic-frame/v1`. |
| `adaptSemanticFrameToGrammarianV1(frame)` | See the semantic-frame section. Failure never mutates the input. |

### Standardized examples

Canonical organic-deconstruct corpus: [`src/catalog/standard-examples.json`](./src/catalog/standard-examples.json)
(`schema_version`: `standard-examples/v1`). There is one file. Package subpath
`oq-api/examples/standard-examples.json` re-exports it. Pages still publishes
`/examples/standard-examples.json` and `/api/v{version}/standard-examples.json`
from that file at publish time.
**Not** an attested-phrase list — see [`docs/standard-examples.md`](./docs/standard-examples.md).

| Export | Contract |
|---|---|
| `STANDARD_EXAMPLES` | Immutable `{ worked, sentences }`. Worked rows require `surface` (single-word showcase). `gloss` is an English `string` or `{ en?, da?, kl? }`. Optional `id`, `chain`, `tags`, `provenance`. `sentences` is reserved/empty in v1 (no attested phrase lists, no gold builds / lattice dumps). |
| `STANDARD_EXAMPLES_SCHEMA` / `STANDARD_EXAMPLES_ID` | `standard-examples/v1` and `oq-api/standard-examples`. |
| `getStandardExamples()` | Defensive copy for bl.oq.gl and other consumers. |
| `citedStandardExamples()` | Subset with provenance or `cited` tag (semantic-case catalog rows). |
| `glossText` / `glossLocales` / `exampleSurface` | Normalize gloss locales and surface from worked/sentence rows. |
| `modalBatch` / `modalExampleBatches` / `toModalExampleItem` | Lab helpers; Misiliineq Simple/Amounts remain lab-local modules. |
| Package export `oq-api/standard-examples.json` | Same JSON bytes as `src/catalog/standard-examples.json`. `oq-api/examples/standard-examples.json` is that same file. |


### Engine options

Optional per-call bag on `buildWord`, `analyzeWord`, and the conjugation
label helpers. Defaults need no browser and do not read `localStorage`.

| Field | Default | Meaning |
|---|---|---|
| `linguistTerms` | `false` | Technical grammar terms instead of plain glosses. |
| `pronounPreference` | `"he"` | For 3sg/4sg labels: `"he"` / `"she"` / `"it"` / `"all"`. Gloss helpers that already took this argument keep their own default (`"all"`). |
| `subjectPronounPreference` | falls back to `pronounPreference` | Subject-role placeholder preference for gloss helpers. |
| `actorPronounPreference` | falls back to `pronounPreference` | Actor-role placeholder preference for gloss helpers. |
| `storage` | omitted | Optional `StorageLike` (`getItem` / `setItem` / `removeItem`) for callers that want to persist engine markers. The library never opens `localStorage` itself. |
| `locale` | active locale (`"en"`) | Morphology-catalog language for this call. |

`EngineOptions` is a TypeScript type in `src/public-api.d.ts`, not a runtime export.

### Conjugation labels — resolved, ready-to-display text

The same friendly labels oq's own "conjugate to…" modal shows for a verb-mood
paradigm coordinate — resolved text, not a raw i18n key or a bare grammar
term. Plain-language by default (`linguistTerms` OFF), falling back to the
technical grammar term for a mood with no plain gloss. `resolvePersonLabel`
additionally honors `pronounPreference` for the two gendered person/number
combinations (3sg/4sg), unless you opt out.

| Export | Contract |
|---|---|
| `resolveMoodLabel(mood, options?)` | `mood` is one of the structured mood keys grammarian's `inflection.mood` publishes, uppercased with any `_DIFF`/`_SAME`/`_TR_GI`/`_DUAL` variant suffix `buildEndingCatalog`'s entries carry (e.g. `"IND"`, `"CAU_DIFF"`) — see `src/engine/conjugation.js`'s `parseEndingEntry`. Returns `{ text, title }`: `text` is the label to show; `title` is the technical term to show as a tooltip when `text` is a plain gloss, or `null` when `text` already IS the technical term (nothing to disclose). |
| `resolvePersonLabel(person, number, opts?)` | `person` a number 1-4, `number` `"SG"`/`"PL"` (case-insensitive). Returns a plain string ("I", "you", "he", …). `opts` may include `EngineOptions` plus `{ ignorePronounPreference? }` — pass `true` to always get the combined "he/she/it" form regardless of `pronounPreference`. |
| `resolveFieldLabel(key, options?)` | `key` one of `"mood"`, `"person"`, `"subject"`, `"object"` — a field heading, not a value. Returns `{ text, title }`, same shape as `resolveMoodLabel`. |
| `t(key, values?, locale?)` | Direct lookup. Pass `locale` (`"en"` / `"da"` / `"kl"`). Omitting it uses English. An unavailable catalog falls back to English. Missing keys return the key itself. GUI chrome catalogs are not shipped. Prefer the resolve\* functions above for conjugation labels. |

### Conjugation math — stem derivation, catalog, conjugate

The same functions oq's conjugation widget and Deconstruct reverse-lookup
use. Additive endings only; undocumented citation patterns return `ok: false`
rather than a guessed stem. Prefer `conjugateForm` when a consumer wants the
same headword→cell→surface path oq's conjugation widget uses; the lower-level
exports remain for catalog UI and Deconstruct. `analyzeWord` is unchanged.

| Export | Contract |
|---|---|
| `parseEndingEntry(entry)` | One grammarian morpheme → conjugation ending, or `null` if it is not a supported additive verb-mood ending. |
| `buildEndingCatalog(flat)` | Parse a grammarian `flat` array into sorted `ConjugationEnding[]`. Non-array ⇒ `[]`. |
| `catalogMoods(catalog)` | Distinct mood keys in catalog order. |
| `endingsForMood(catalog, mood)` | Endings belonging to one mood. |
| `subjectsForMood(catalog, mood)` | Distinct transitive subjects for a mood; empty for intransitive. |
| `objectsForSubject(catalog, mood, person, number)` | Distinct objects for one subject in a transitive mood. |
| `findEnding(catalog, mood, subjPerson, subjNumber, objPerson?, objNumber?)` | One cell, or `null`. Omit object args (or pass null) for an intransitive cell. |
| `matchConjugationEnding(word, catalog)` | Longest unique catalog ending that is a suffix of `word`. `{ ok, stem, ending }`; declines on no match or equal-length ambiguity. |
| `deriveIntransitiveStem(headword, stemHint?)` | Strip documented 3sg `-voq`/`-poq`. `stemHint` trusted outright (Word Builder / Deconstruct / katersat); doubled-consonant remainders decline without it. `{ ok, stem }`. |
| `deriveTransitiveStem(headword, stemHint?)` | Strip documented 3sg/3sg `-aa`/`-vaa`. `stemHint` from katersat is trusted; doubled-consonant remainders decline without it. |
| `deriveSchwaStem(headword)` | Strip documented schwa-stem `-qaaq`. |
| `canConjugate(headword)` | True if any of the three stem derivations succeeds. |
| `conjugate(stem, ending)` | `{ word, approximate }`. `ending` needs `{ id, text }`. Approximate when `OUT_OF_SCOPE_IDS` contains the ending. |
| `conjugateSchwaStem(stem, ending)` | Uses published `allomorphs.schwa_final_stem`. `{ ok: true, word, approximate: true }` or `{ ok: false, word: null }`. |
| `conjugateForm(headword, spec)` | One-call transform: citation headword + paradigm cell → surface. `spec` needs `catalog` (from `buildEndingCatalog`), `mood`, and `subject` `{ person, number }` (or top-level `person`/`number`); optional `object` (implies transitive), `transitive`, `stemHint` (honored for **intransitive and transitive**), `gloss`. Returns `{ ok, word, approximate, stem, ending, schwa, errorKey, translation }`. Stable `errorKey`s: `empty_catalog`, `invalid_selection`, `missing_ending`, `unsupported_stem`, `schwa_unsupported`. No interrogative `?` / syllabify — presentation stays in the consumer. Never invents a stem. |
| `haveNForm(spec)` | One-call have-N transform: host noun + optional numeral/quantifier/howMany + `-qaq` (`N_qaq_Vb`) + conjugation ending → surface(s) via **`buildWord` seq** (never invented stems). `spec` needs `host` (id / raw entry / preset) and either `ending` or `endingCatalog`+`mood`+`subject`/`person`/`number`; optional `catalog` (by-id map or list), `count` (1–10 or cardinal id) / `quantifier` / `howMany: true` (→ `HAVE_N_HOW_MANY` = `qassit`, #780), `many: true` (→ `HAVE_N_MANY` = `qassiit`, #781), `truth: true` / `"ilumut?"` (→ prepose pedagogy `ilumut?`, grammarian #783) or `truth: "ilumut"` (assertive), `intensifier` / `affixes` (between `-qaq` and ending; `HAVE_N_INTENSIFIERS` = `V_luinnaq_Vb` / `V_vik_Vb`), `qaq` (default `N_qaq_Vb`). Pair `howMany` with interrogative mood (e.g. INTERR 2SG) for “how many N do you have?”. Truth wraps combine with many/howMany/count/intensifiers for classics like “is it true that you have many dogs?”. Returns `{ ok, word, approximate, closed, seq, host, qaq, ending, numeral, truth, phrase, errorKey, missingIds }`. Numeral/quantifier (when requested) is a separate instrumental companion (`N_INS_SG` for one, `N_INS_PL` otherwise; Bjørnum K6§2 INS via `numeralInstrumentalSurface` (`ataatsimik`, `marlunnik`, `pingasunik`, `sisamanik`, `tallimanik`, `arfinilinnik`, `arfineq-marlunnik`, `arfineq-pingasunik`, `qulinik`, `qassinik`, `qassiinik`)). Stable `errorKey`s: `missing_host`, `missing_qaq`, `missing_ending`, `missing_numeral`, `missing_affix`, `missing_particle`, `invalid_selection`, `empty_catalog`, `build_failed`. `HAVE_N_CARDINALS` maps 1–10 → catalog ids (`arfinillit` for quantity-six; #779 fills 3/4/7/8; #785). `HAVE_N_TRUTH` = `ilumut`; `HAVE_N_TRUTH_SURFACES` = `{ assertive: "ilumut", interrogative: "ilumut?" }`. |
| `glossSentence(meaning, gloss, person?, number?)` | English-only: fill the ending's trailing `(...V...)` clause with the verb gloss. `null` when unparseable. |
| `moodSubjectSplitKey(mood)` | i18n key for same/different-subject hint (`CAU_SAME` / `CAU_DIFF`), or `null`. |

### Dictionary lookup

Network-backed: these fetch full upstream dictionary JSON on first use and
cache in memory afterwards.

| Export | Contract |
|---|---|
| `DICT_SOURCES`, `KAT_SOURCES` | Source registries to pass to `searchEntries`. |
| `searchEntries(sources, query, options?)` | Async search: `{ results, attributions, failed, errors }`. Options: `{ lang?, match?, rank?, advancedRegex? }`; result entries also expose `surface`, `wordClass`, and `glosses`. Rank/match live here (not in oq's dictionary view). |
| `findExactDictMatch(word)` | Async exact Kalaallisut-headword lookup; first matching entry or `null`. Trailing citation punctuation (`tuavi!`) is not part of the lexeme. Never throws. |
| `mergeBySurface(entries)` | Homograph merge: collapse same-headword+class entries, unioning glosses. Pure; used by `getAllEntries`. |
| `getAllEntries(sources)` | Cached full-dictionary dump for the given `DICT_SOURCES` (or a subset). Empty until those sources have been loaded. Cloned for the caller. |
| `loadFullSource(src)` | Fetch one `DICT_SOURCES` / `KAT_SOURCES` entry into the in-memory cache. Deduplicates concurrent calls; no-ops if already loaded. Call this before `getAllEntries` / `injectDictionaryRoots` if you have not searched yet. `searchEntries` and `findExactDictMatch` load on their own. |
| `ensureFullLoaded(sources)` | `Promise.allSettled` of `loadFullSource` over `sources`. |
| `randomEntry()` | Random `DICT_SOURCES` entry after loading the full lexicon, or `null`. Cloned. |
| `injectDictionaryRoots(entries, presets?, opts?)` | Synthetic stem presets (`id: dict_…`) from dictionary headwords. No fetch. Skips non-N/V POS except interjections / bang-lemmas (`tuavi!` → sealed `tuavi`). Trailing `!`/`?` is stored as `citationPunct` and reattached on display (card / sentence KL surface / gloss if the translation had none). Engine `expected` stays unpunctuated so rebuild can match typed `tuavi`. By default skips stems already in `presets`. `noun_plural_form`, finite verbs, exclamations, and **citation abbreviations** (gloss `abbrev.` / `forkort…`, or upstream `forkortes` metadata; oq-grammarian#795) get `continuation_class: WORD_FINAL` on **exact** citations. Proper-prefix injection of citation abbreviations is refused (`m` ⊏ `maanna`). Optional `opts.surfaces` scopes injection to exact hits and (unless `includePrefixes: false`) the longest proper-prefix stem per surface — prefix injections stay unsealed so morphology can attach (api.oq.gl#200). Returns only the new presets — concat onto the catalog yourself. Does not change `analyzeWord`. `analyzeSentence(..., { dictEntries })` applies the same injection for the sentence token surfaces. |
| `suggestWordCompletions(prefix, maxCount?)` | Async dictionary prefix completions (excludes the exact prefix). `{ word, gloss, entry? }[]`. Distinct from `suggestMorphCompletions`. |

### Classification catalogs and katersat hints

Network-backed like dictionary lookup: live fetch on first use (katersat has
no vendored `data-mirror` — GPL-3.0-or-later, live-fetch only). Pass
`options.data` on the loaders to skip the network with caller-supplied JSON.
oq's PWA mirror stays an oq concern.

| Export | Contract |
|---|---|
| `SEMANTIC_CLASSES_URL` | Live katersat `semantic_classes.json` URL. |
| `loadSemanticClasses(options?)` | Warm the catalog. `options.data` is the raw JSON document. Failures resolve to `null`. |
| `getSemanticClasses()` | Cached class list, or `[]` before a successful load. |
| `getSemanticClassByCode(code)` | One class or `null`. |
| `getSemanticClassById(id)` | One class or `null`. |
| `getSemanticClassChildren(id)` | Direct children of `id`, or `[]`. |
| `DOMAINS_URL` | Live katersat `domains.json` URL. |
| `loadDomains(options?)` | Warm the domain catalog. `options.data` is the raw JSON. Failures resolve to `null`. |
| `getDomains()` | Cached domains, or `[]`. |
| `getDomainByCode(code)` | One domain or `null`. |
| `getKatersatByLetterUrl(letter)` | URL of one katersat `by-letter/<letter>.json` shard. |
| `getKatersatTransitiveStemHint(headword, options?)` | Bare stem from katersat `fst_analyses`, or `null`. Live-fetches one letter shard; never throws. |
| `getKatersatGlossHint(headword, options?)` | First published English verb gloss from katersat, or `null`. Live-fetches one letter shard; never throws. |

## Poking at upstream data directly

You don't need `oq` or Node to look at the raw data — the upstream sources
(listed in `src/catalog/upstream-sources.js`) are plain JSON files served over
HTTPS, so `curl` + `jq` (and `sqlite3` for anything you want to query
repeatedly) get you there directly. This is not an API `oq` exposes — it's
just how to inspect the same files `oq` fetches.

### Oqaasileriffik dictionary — `curl` + `jq`

Full entry shape: `{id, lexeme, word_class, class_path, stem, gloss_en,
source_file, source_row}`.

```bash
# meta + how many entries
curl -s https://jandahl.github.io/Oqaasileriffik-dicts/all_entries.json \
  | jq '{meta: .meta.attribution, count: (.dictionary_entries | length)}'

# grep for a headword substring, projected to just the fields you probably want
curl -s https://jandahl.github.io/Oqaasileriffik-dicts/all_entries.json \
  | jq '.dictionary_entries[]
        | select(.lexeme | test("qimme"))
        | {lexeme, word_class, gloss_en}'
```

### Katersat lexicon — sharded by letter, so fetch only what you need

Full lexeme shape includes `id, kalaallisut, english, danish, word_class,
semantic_classes, valence, domain, gender, fst_analyses, definition, info,
verb_frames, …` — most of it null/empty for a given entry, so projecting
down to what you actually need keeps output readable.

```bash
# one letter's shard instead of the whole lexicon, just the useful fields
curl -s https://jandahl.github.io/Oqaasileriffik-katersat/by-letter/q.json \
  | jq '.lexemes[] | {kalaallisut, word_class, danish, english}'

# every verb ("v") in a shard, headword + Danish gloss only
curl -s https://jandahl.github.io/Oqaasileriffik-katersat/by-letter/q.json \
  | jq '[.lexemes[] | select(.word_class == "v") | {kalaallisut, danish}]'
```

### Grammarian morphemes (version-pinned) — `curl` + `jq`

Full entry shape (under `.by_id` on the current `GRAMMAR_MORPHEMES_URL`, or
legacy `.flat[]`) nests fields under `lexical_facts`, `application_logic`,
`plain_gloss`, etc. — project the ones you want rather than reading the
whole nested object. JS consumers should call `flattenGrammarianData(data)`
rather than assuming `.flat`.

```bash
curl -s https://grammarian.oq.gl/v2/grammar/morphemes-by-id.json \
  | jq '.meta'

# id, underlying form, English gloss, and category for every morpheme
curl -s https://grammarian.oq.gl/v2/grammar/morphemes-by-id.json \
  | jq '[.by_id[]] | flatten | .[] | {
      id,
      underlying_form: .application_logic.underlying_form,
      gloss: .plain_gloss.en,
      category
    }'
```

### Loading a source into SQLite for repeated querying

Fetching + `jq`-filtering on every query gets slow once you're doing more
than a couple of lookups. `sqlite3`'s JSON extension (built in on recent
versions) can load a fetched file once and let you query it with SQL:

```bash
curl -s https://jandahl.github.io/Oqaasileriffik-dicts/all_entries.json \
  -o /tmp/all_entries.json

sqlite3 :memory: <<'SQL'
.mode json
CREATE TABLE entries AS
  SELECT value ->> 'id'         AS id,
         value ->> 'lexeme'     AS lexeme,
         value ->> 'word_class' AS word_class,
         value ->> 'gloss_en'   AS gloss_en
  FROM json_each(readfile('/tmp/all_entries.json'), '$.dictionary_entries');

SELECT id, lexeme, gloss_en
FROM entries
WHERE lexeme LIKE 'qimme%'
ORDER BY lexeme;
SQL
```

Same pattern works for the katersat lexicon (`$.lexemes` instead of
`$.dictionary_entries`) — swap in whichever source URL and JSON path you
need. `docs/data-mirror/*.json` (see `src/catalog/upstream-sources.js`'s `local`
field) are committed same-origin snapshots of some of these sources, so once
`oq` is deployed you can `curl` those instead if you want a pinned,
build-time copy rather than live upstream data.

## Changing the surface

1. Edit the export list in `src/public-api.js` and bump `version` in `package.json` (`API_VERSION` reads it).
2. Update the tables above.
3. Update `EXPECTED_SURFACE` in `test/public-api.test.js`.

The snapshot test fails until all three agree — that friction is the point.
