Skip to content

IPA: strong English forms, standard accents, Icelandic stress, Chinese readings - #221

Merged
xuelink merged 1 commit into
mainfrom
claude/tools-ipa-polish
Sep 24, 2026
Merged

xuelink merged 1 commit into
mainfrom
claude/tools-ipa-polish

Conversation

@xuelink

@xuelink xuelink commented Sep 24, 2026

Copy link
Copy Markdown
Member

Fixes the three known imperfections left after #219, and a hidden filter bug that turned up while doing it.

1. English: citation forms, not weak forms

Before, "it" showed /ət/, "in" /ən/, "that" /ðət/ and "you" /jɪ/. Wiktionary lists the weak form first. Two causes:

  • "it": the strong /ɪt/ is tagged General American, RP and Australian. The regional filter threw the whole transcription out because of "Australia".
  • "in", "that": untagged; the weak form simply came first.

For English, a reading with a full vowel now wins over one with only a reduced vowel. This also outweighs the US preference, because for "you" only the British transcription is strong.

Result: it /ɪt/, in /ɪn/, that /ðæt/, you /ju/, have /hæv/, could /kʊd/.

2. Regional filter, split in two (affects every language)

  • Dated or dialectal tags (obsolete, dialectal, Tiberian, Classical…) always drop a transcription.
  • Region names (Australia, Taiwan, Quebec, Scotland…) drop it only if the page's own accent isn't also tagged.

The old single filter dropped too much, and dialect readings slipped through in the gaps. Examples of what changes:

Language Word Before After
Chinese 告訴 /kau²¹³ su²¹³/ (Sichuanese-style tones) /kɑʊ̯⁵¹ su¹/
Chinese 哥哥 /kə²⁴ kə²⁴/ /kɤ⁵⁵ g̊ə²/
French merci /maɛ̯ʁ.si/ (Quebec) /mɛʁ.si/
French tant /tã/ /tɑ̃/
English now /nuː/ (Scottish) /naʊ/
English good /ɡɵd/ /ɡʊd/
English dead /diːd/ /dɛd/

3. Icelandic rows from goruut

  • They now get the first-syllable stress mark, as Icelandic stress always falls there.
  • A reading with fewer syllables than the spelling is dropped. This removes 13 rows; "skenkja" was /skrmkja/. The check is Icelandic-only: in Hungarian ("gy", "ny") and Ukrainian (і, ї, є), vowel letters don't map one-to-one onto syllables.

4. Chinese characters with several readings

A single character now takes the reading the list's own words use most, weighted by how common they are. For example, 還 appears as hái in 還是 and 還有, so it gets /xaɪ̯³⁵/ rather than huán.

Changed rows

Language Rows changed
Icelandic 452 (mostly the added stress mark)
Chinese 195
English 180
French 12
Dutch 1
Spanish 1

I spot-checked the diffs. They are corrections or neutral variants; the one debatable change is French "juin", /ʒɥɛ̃/ → /ʒwɛ̃/. Empty rows go from 481 to 494, because of the 13 Icelandic rows dropped.

Checked: svelte-check shows 0 errors. Its 8 warnings are all in files this PR doesn't touch. The rebase onto main was clean.

🤖 Generated with Claude Code

…e readings

- English shows a function word's citation form, not its weak one: "it"
  /ɪt/ not /ət/, "in" /ɪn/, "that" /ðæt/, "you" /ju/.
- The filter for dated and regional pronunciations is split in two: a
  dated or dialectal tag always drops a transcription, but a region's name
  drops it only if the page's own accent is not tagged too. /ɪt/ is tagged
  General American, RP and Australian at once, and Standard Mandarin
  transcriptions also tagged "Taiwan" were being dropped — which let
  Mandarin dialect readings through (告訴 /kau²¹³ su²¹³/, now /kɑʊ̯⁵¹ su¹/)
  and Quebec French into the French list (merci /maɛ̯ʁ.si/, now /mɛʁ.si/),
  and in English a Scottish "now" /nuː/, now /naʊ/.
- Icelandic readings from goruut get the first-syllable stress mark, and a
  reading with fewer syllables than the spelling is dropped (skenkja had
  been /skrmkja/): 13 rows.
- A Chinese character with several readings takes the one the list's own
  words use most: 還 /xaɪ̯³⁵/ (hái, as in 還是), not huán.

Changed rows: Chinese 195, English 180, Icelandic 452 (mostly the stress
mark), French 12, Dutch 1, Spanish 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@xuelink
xuelink merged commit bb4c733 into main Sep 24, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant