Skip to content
LontaraBahasa Indonesia

Ejaan

The Latin side needs as much specification as the Lontara side. An underspecified Latin orthography makes the whole enumeration unsound.

rules.json v0.2.0 · unreviewed · 11 rules

What is specified here

Latin onsets are taken from each letter’s Unicode character name. The inherent vowel /a/ has no sign.
OnsetLetterUnicode name
kBUGINESE LETTER KA
gBUGINESE LETTER GA
ngBUGINESE LETTER NGA
ngkprenasalBUGINESE LETTER NGKA
pBUGINESE LETTER PA
bBUGINESE LETTER BA
mBUGINESE LETTER MA
mpprenasalBUGINESE LETTER MPA
tBUGINESE LETTER TA
dBUGINESE LETTER DA
nBUGINESE LETTER NA
nrprenasalBUGINESE LETTER NRA
cBUGINESE LETTER CA
jBUGINESE LETTER JA
nyBUGINESE LETTER NYA
nycprenasalBUGINESE LETTER NYCA
yBUGINESE LETTER YA
rBUGINESE LETTER RA
lBUGINESE LETTER LA
wBUGINESE LETTER VA
sBUGINESE LETTER SA
BUGINESE LETTER A
hBUGINESE LETTER HA
Each vowel sign, its codepoint, and the basis for the Latin value used here.
VowelSignBasis
ano sign — the inherent vowelEvery consonant letter carries the inherent vowel /a/. A vowel sign replaces it; there is no way to remove it.
iU+1A17 · BUGINESE VOWEL SIGN IUCD canonical combining class 230 (Above)
uU+1A18 · BUGINESE VOWEL SIGN UUCD canonical combining class 220 (Below)
éProvisional — awaiting a reviewerU+1A19 · BUGINESE VOWEL SIGN ELatin `é`, NOT `e`. Taken from attested pairs in data/corpus/bugwiki-pairs.json rather than from the Unicode character name: Sapéda, Sulawési, Baruné, Watamponé, Sawérigading, Éropa, Lébarepuluq and Komputeréq all write `é` with this sign. PROVISIONAL — Wikipedia is community practice, not an authority. See openQuestions.vowel-sign-e-ae.
oU+1A1A · BUGINESE VOWEL SIGN OSpacing combining mark; the UCD assigns no combining class. See the rendering conformance page.
eProvisional — awaiting a reviewerU+1A1B · BUGINESE VOWEL SIGN AELatin `e`, NOT `ae`. `ae` came from the Unicode character name and was wrong: about thirty independent attested pairs in data/corpus/bugwiki-pairs.json write plain `e` with this sign (Bekasi, Wettu, Peddang, Jeppang, Halaleng, Balangeng, Irelang, Banjaramaseng and more). PROVISIONAL — Wikipedia is community practice, not an authority. See openQuestions.vowel-sign-e-ae.

The Latin distinctions the four classes are defined against

The Latin side has to encode final consonants, gemination, prenasalisation and the glottal stop — precisely the distinctions the ambiguity classes are defined against. If the Latin side is vague, the whole enumeration is unsound.

final consonant

A syllable-final consonant is not written.

lontara.final.drop · derived

PRD §2; inventory.json block.hasVirama = false

gemination

A single consonant and a doubled consonant are written alike.

lontara.gemination.collapse · derived

PRD §2; inventory.json block.hasVirama = false

prenasalisation

Prenasalisation is not written where no prenasal letter covers the cluster.

lontara.prenasal.drop · derived

PRD §2; inventory.json block.hasVirama = false

glottal stop

The glottal stop is not written.

lontara.glottal.drop · derived

PRD §2; inventory.json block.hasVirama = false

Where practice diverges

  • latin.glottal.qprovisional

    Accept `q` as an input spelling of the glottal stop.

    Attested pairs in data/corpus/bugwiki-pairs.json (Bugis Wikipedia, pages-articles dump 20260701, sha256 37b7da2c…, CC BY-SA 4.0) write a word-final `q` that is not written in Lontara at all: Kamisiq → ᨀᨆᨗᨔᨗ, Goloq → ᨁᨚᨒᨚ, Ahéraq → ᨕᨖᨙᨑ, Perancisiq, Donatturampeq, Lébarepuluq, Komputeréq. PRD §6.7 names the glottal stop as what Bugis Latin writes with an apostrophe, so `q` is read as a second convention for the same phoneme and normalised to U+0027.

    PROVISIONAL. That `q` here is the glottal stop rather than a final /k/ is an inference: both are unwritten in Lontara, so THE OUTPUT IS IDENTICAL EITHER WAY and only the declared ambiguity class differs — `glottal` versus `final`. The risk of being wrong is therefore confined to a label, which is why this ships rather than waiting. A reviewer should still settle it. See openQuestions.glottal-q.

  • latin.wa.vprovisional

    Accept `v` as input for the onset written by U+1A13, whose Latin form is `w`.

    Attested pairs in data/corpus/bugwiki-pairs.json (Bugis Wikipedia, pages-articles dump 20260701, sha256 37b7da2c…, CC BY-SA 4.0) write U+1A13 as Latin `w` in every case (Watangpola, Tawawu, Kuweng, Awayeng, Sulawési, Noruwégiya, Réiwa, Pérétiwi); `v` never appears. inventory.json therefore takes `w` as the onset, and this rule accepts `v` on input for the same letter, since the Unicode character name BUGINESE LETTER VA leads people to type it.

    PROVISIONAL. Bugis Wikipedia is community practice, not an authority, and the direction of this rule was inverted in v0.2.0 on the strength of that corpus — it previously mapped `w` to `v` on the basis of the Unicode character name alone. It only widens what is accepted as input; it never changes what is written out. A reviewer should confirm `w` is standard. See openQuestions.va-latin.

Open questions

Not guessed at. Each one names who to ask — and how many forms its answer decides.

Ordered by how many forms each affects, out of the 1,321 distinct forms in the lexicon. The number says which forms depend on the answer — not what they would become. Knowing that requires the answer first.

  1. openQuestions.final-inventory

    569 of 1,321 forms

    Which consonants can close a syllable in Bugis, and with what relative frequency?

    It would let the reader enumerate `final` readings structurally rather than only where the lexicon attests them. Deliberately NOT guessed at: a structural enumerator seeded with an invented set of possible finals would produce a confident-looking tree of readings that are not Bugis. Until the set is cited, enumeration stays lexicon-driven and says so.

    For example 'uang · abad · abbatireng · aceh · adalah · adan · addatuang · administrasi … and 561 more

    The reader cannot branch structurally over syllable-final consonants without a cited inventory of them. Every form the writer records a final loss on is a form that inventory would have to account for.

    This counts which forms depend on the answer, not what they would become. Running the alternative needs the very rule that is not settled — so the size is computable and the consequence is not.

    Ask: The dictionaries in PRD Appendix A; a Bugis reviewer

  2. openQuestions.vowel-sign-e-ae

    565 of 1,321 forms

    Which Latin vowels do U+1A19 VOWEL SIGN E and U+1A1B VOWEL SIGN AE correspond to?

    Answered in practice and CHANGED in v0.2.0. The Unicode character names gave `e` for U+1A19 and `ae` for U+1A1B; roughly thirty independent attested pairs say the opposite assignment — plain `e` takes U+1A1B (Bekasi, Wettu, Peddang, Jeppang, Halaleng) and `é` takes U+1A19 (Sapéda, Sulawési, Baruné, Watamponé, Sawérigading). The old mapping meant the writer silently dropped every `é`, losing a whole syllable in a word like Sapéda. What remains open is whether this is standard orthography, and what phonemes the two signs represent — /ə/ versus /e/ is the obvious guess and a guess is all it is.

    For example abbatireng · abbatirenna · aceh · ade · ade' · ade'na · agnetha · aheli … and 557 more

    The question is which Latin vowel each of these two signs carries, so every form written with either sign is a form whose Latin reading the answer decides. v0.2.0 already swapped them once; this is the set that moved.

    This counts which forms depend on the answer, not what they would become. Running the alternative needs the very rule that is not settled — so the size is computable and the consequence is not.

    Ask: Badan Bahasa South Sulawesi; a Bugis reviewer

  3. openQuestions.vowel-sign-position

    381 of 1,321 forms

    Which side do VOWEL SIGN E (U+1A19) and VOWEL SIGN O (U+1A1A) render on?

    Both are spacing marks (Mc) with no UCD combining class, so the encoding does not answer it. This is a rendering question, and the conformance page answers it by eye rather than inventory.json asserting it.

    For example akkarungengngé · aléna · alérapanna · aloagen · ambo · ambo'na · amérika · andorra … and 373 more

    A rendering question rather than an orthographic one, so the affected set is every form written with either sign — those are the forms that would look wrong on a device that places them on the other side.

    This counts which forms depend on the answer, not what they would become. Running the alternative needs the very rule that is not settled — so the size is computable and the consequence is not.

    Ask: Nobody — check /aksara/konformansi on a real device and record the answer.

  4. openQuestions.va-latin

    112 of 1,321 forms

    Is the Latin form of U+1A13 `w` (as every attested pair writes it) or `v` (as the Unicode character name suggests)?

    Settled in practice, not in principle. Every one of the attested pairs in data/corpus/bugwiki-pairs.json writes `w` and none writes `v`, so v0.2.0 takes `w` as the onset and accepts `v` only on input. What is still missing is confirmation that `w` is standard rather than merely what Wikipedia contributors do.

    For example aruwa · awa · battowa · battowaé · bawa · bawang · bawi · dewata … and 104 more

    U+1A13 is the only letter whose Latin onset this question disputes, so the forms whose written output contains it are exactly the forms whose reading changes if the answer is `v` rather than `w`.

    This counts which forms depend on the answer, not what they would become. Running the alternative needs the very rule that is not settled — so the size is computable and the consequence is not.

    Ask: Badan Bahasa South Sulawesi; the dictionaries in PRD Appendix A

    Blocks: latin.wa.v

  5. openQuestions.glottal-q

    13 of 1,321 forms

    Is word-final `q` in Bugis Latin the glottal stop, or a final /k/, or does it vary?

    It changes which ambiguity class the writer declares, and nothing else — both are unwritten in Lontara, so the aksara output is byte-identical either way. Recorded so the label can be corrected rather than quietly trusted.

    For example 'uang · aba' · bentu' · eppa' · iyaré'ga · kappala' · kappala'é · kappala'na … and 5 more

    The question is what `q` in Bugis Latin represents, so the affected forms are the ones the lexicon actually spells with it. Matched anywhere in the form rather than word-finally only, which over-selects rather than under-selects — the wider set is the honest one to hand a reviewer.

    This counts which forms depend on the answer, not what they would become. Running the alternative needs the very rule that is not settled — so the size is computable and the consequence is not.

    Ask: Badan Bahasa South Sulawesi; a Bugis reviewer

    Blocks: latin.glottal.q

  6. openQuestions.prenasal-coverage

    8 of 1,321 forms

    Four prenasal clusters are written with dedicated letters (ngk, mp, nr, nyc). What happens to every other nasal+stop cluster?

    PRD §2 states flatly that prenasalisation is not written, but the block plainly writes four of them. Either the statement is a simplification or the four letters are used differently than their names suggest. The reader's enumeration of the `prenasal` class depends on the answer. TWO attested pairs bear on it, and both point the same way — that the prenasal letter is chosen by the SECOND element, whichever nasal precedes it: Watangpola → ᨓᨈᨇᨚᨒ writes `ngp` with U+1A07 MPA, and Perancisiq → ᨄᨛᨑᨏᨗᨔᨗ writes `nc` with U+1A0F NYCA. DELIBERATELY NOT ACTED ON. Two instances is thin, either could be an article error, and generalising would change `manka` from MA KA to MA NGKA — a real behaviour change on weak evidence. The rule set still drops such a nasal and declares the loss, and these two pairs are the reason to ask a reviewer rather than the reason to guess.

    For example ancajingeng · bencana · diluncurkan · mancaaji · mancaji · perancis · perancisi · ripancaji

    A nasal that the writer drops and declares as the prenasal class is precisely a nasal the four dedicated letters did not cover. If the letters generalise, these are the forms that stop losing it.

    This counts which forms depend on the answer, not what they would become. Running the alternative needs the very rule that is not settled — so the size is computable and the consequence is not.

    Ask: A Bugis reviewer who reads manuscript Lontara; Everson's encoding proposals (PRD Appendix A)

    Blocks: lontara.prenasal.drop

  7. openQuestions.pallawa

    0 of 1,321 forms

    What does PALLAWA (U+1A1E) separate, and how should it and END OF SECTION (U+1A1F) be represented in Latin?

    The writer cannot emit either mark and the reader cannot read one back until this is answered. inventory.json carries `latin: null` for both rather than assuming pallawa is a space.

    The two punctuation marks are the whole subject of the question. A count of zero over a single-word lexicon is itself the finding: this question cannot be sized until the repository holds running text.

    Ask: The dictionaries and textbook sources in PRD Appendix A; a Bugis reviewer

Audit it yourself

No orthographic rule is written in application code. All of it is in data/rules/rules.json with an id, a priority and a citation — so that a Bugis-literate reviewer who does not program can audit it. `pnpm rules:report` prints it as a readable table.

Aksara