IPA: strong English forms, standard accents, Icelandic stress, Chinese readings - #221
Merged
Merged
Conversation
…e readings - English shows a function word's citation form, not its weak one: "it" /ɪt/ not /ət/, "in" /ɪn/, "that" /ðæt/, "you" /ju/. - The filter for dated and regional pronunciations is split in two: a dated or dialectal tag always drops a transcription, but a region's name drops it only if the page's own accent is not tagged too. /ɪt/ is tagged General American, RP and Australian at once, and Standard Mandarin transcriptions also tagged "Taiwan" were being dropped — which let Mandarin dialect readings through (告訴 /kau²¹³ su²¹³/, now /kɑʊ̯⁵¹ su¹/) and Quebec French into the French list (merci /maɛ̯ʁ.si/, now /mɛʁ.si/), and in English a Scottish "now" /nuː/, now /naʊ/. - Icelandic readings from goruut get the first-syllable stress mark, and a reading with fewer syllables than the spelling is dropped (skenkja had been /skrmkja/): 13 rows. - A Chinese character with several readings takes the one the list's own words use most: 還 /xaɪ̯³⁵/ (hái, as in 還是), not huán. Changed rows: Chinese 195, English 180, Icelandic 452 (mostly the stress mark), French 12, Dutch 1, Spanish 1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the three known imperfections left after #219, and a hidden filter bug that turned up while doing it.
1. English: citation forms, not weak forms
Before, "it" showed /ət/, "in" /ən/, "that" /ðət/ and "you" /jɪ/. Wiktionary lists the weak form first. Two causes:
For English, a reading with a full vowel now wins over one with only a reduced vowel. This also outweighs the US preference, because for "you" only the British transcription is strong.
Result: it /ɪt/, in /ɪn/, that /ðæt/, you /ju/, have /hæv/, could /kʊd/.
2. Regional filter, split in two (affects every language)
The old single filter dropped too much, and dialect readings slipped through in the gaps. Examples of what changes:
3. Icelandic rows from goruut
4. Chinese characters with several readings
A single character now takes the reading the list's own words use most, weighted by how common they are. For example, 還 appears as hái in 還是 and 還有, so it gets /xaɪ̯³⁵/ rather than huán.
Changed rows
I spot-checked the diffs. They are corrections or neutral variants; the one debatable change is French "juin", /ʒɥɛ̃/ → /ʒwɛ̃/. Empty rows go from 481 to 494, because of the 13 Icelandic rows dropped.
Checked:
svelte-checkshows 0 errors. Its 8 warnings are all in files this PR doesn't touch. The rebase onto main was clean.🤖 Generated with Claude Code