@reg (for the normalized form in e.g. historical corpora)
I see number of problems with introducing this attribute:
- It breaks the TEI dictum that attributes should not contain strings, and while it might be true that
@reg will seldom use <g>, it seems bad practice to break a convention to accomodate one case only.
- More importantly, in historical and CMC data two (or more) tokens can regularise to one standard token or vice versa, which cannot be handled by
@reg.
- Also, it can make a difference if you are assigning a PoS tag (MSD) or lemma to the original word or to its regularised equivalent, e.g. a historical word might have a different gender from its modern form, and the lemma of the original word will be different from that of the regularised one - again, this cannot be accomodated in this proposal.
So, I'd propose to sticking to <choice> with <orig> and <reg> and these then containing <w> etc. Yes, somewhat verbose, but covers all the cases (except for discontinuous elements, but that is a whole dimension of extra complication).
I see number of problems with introducing this attribute:
@regwill seldom use<g>, it seems bad practice to break a convention to accomodate one case only.@reg.So, I'd propose to sticking to
<choice>with<orig>and<reg>and these then containing<w>etc. Yes, somewhat verbose, but covers all the cases (except for discontinuous elements, but that is a whole dimension of extra complication).