Repository navigation
matching always behaves as if "-c" was specified #273
Description
Activity
there's also a similar issue with
%s's precision and field width options:$ goawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' ╋:╋ $ gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' ╋:╋ $ LC_ALL=C gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' [garbled output]curiously enough
%cdoesn't have this problem:$ goawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }' [garbled output] $ gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }' ╋:╋ $ LC_ALL=C gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }' [garbled output]So my understanding was that it works in the C/POSIX locale by default (logically) regardless of the system locale (LC_* values), and using -c switches to some "Unicode mode", where "Unicode mode" probably means UTF-8.
your understanding is correct; to use the system locale as strictly required by POSIX, goawk would have to use C wide characters and the relevant libc functions (which it does not, probably for the best!)
also an extremely pedantic side note
the C/POSIX locale character encoding doesn't have to be ASCII! the list of characters it's required to support encoding and the other requirements are explained here
in practice it is ASCII on almost every system nowadays, though i'm aware of a couple systems that still use EBCDIC for the POSIX locale
- changed the title
[-]matching alway behaves as if "-c" was specified[/-][+]matching always behaves as if "-c" was specified[/+]on Feb 28, 2026 Yes, you're right -- these are inconsistencies in bytes/unicode handling. However, the match one in particular is tricky to fix, because Go's
regexpis always UTF-8 (you can't put it in bytes/ASCII mode). I'll leave this issue open to track.there's also a similar issue with
%s's precision and field width options:$ goawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' ╋:╋ $ gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' ╋:╋Interesting (took me few seconds to realize that the 2nd line is gnu awk...).
So this looks correct in gnu awk - assuming the system locale is some UTF-8, but incorrect in gOawk, because it wants to behave in the C locale without
-c.In Unicode/UTF-8 mode, the
%.1sextracts a single codepoint fully, and the%.3swants to extract 3 codepoints but runs out of string after 1.$ LC_ALL=C gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' [garbled output]For reference, the "garbled output", at least for me, seems correct for the C locale: I get the full 3 bytes of the 1st string, followed by
:, followed by the 1st byte of the 2nd string, and then newline:$ LC_ALL=C gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' | od -c 0000000 342 225 213 : 342 \n 0000006
curiously enough
%cdoesn't have this problem:$ goawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }' [garbled output] $ gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }' ╋:╋ $ LC_ALL=C gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }' [garbled output]Precision for
%chas undefined behavior in both C99 and POSIX 2024, and as far as I can tell it's not augmented in the POSIX awkprintfspec. I think in all 3 examples the precision is ignored, which is a valid behavior.But it is true that in this case gOawk does behave as intended -
%cextracts the first char of the following string, which without-cis expected to be the first byte - which indeed it is.And for completion,
goawk -c 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }'indeed extracts the full codepoint in both arguments (and the precision is ignored again).So bottom line, precision for
%swithout-cis wrong and still counts in codepoints, but%cis correct and extracts the 1st byte/codepoint depending on-c.the C/POSIX locale character encoding doesn't have to be ASCII! the list of characters it's required to support encoding and the other requirements are explained here
Thanks. I'm familiar with this page. For reference, the same page in POSIX 2024.
In what sense do you consider it "doesn't have to be ASCII"? the portable character set is ASCII as far as I can tell. Do you mean that like few control codes which I missed are different from ASCII? or that it can be different entirely (and therefore this system is not POSIX)?
For reference, 7.2 POSIX locale says this:
Conforming systems shall provide a POSIX locale, also known as the C locale. In POSIX.1 the requirements for the POSIX locale are more extensive than the requirements for the C locale as specified in the ISO C standard. However, in a conforming POSIX implementation, the POSIX locale and the C locale are identical.
And that page also specifies the POSIX locale chars, which, again, are ASCII, or at least have identical values as ASCII for the vast majority of chars.
i'm aware of a couple systems that still use EBCDIC for the POSIX locale
Interesting. Are there such systems which can be emulated today? I wouldn't mind doing some tests with C code on such system, to test portability (same for systems where byte is not 8 bits).
the match one in particular is tricky to fix, because Go's
regexpis always UTF-8 (you can't put it in bytes/ASCII mode). I'll leave this issue open to track.Yeah, I can understand that. Still worth documenting it someplace, probably at the README. Both the fact that by default it tries to behave in bytes mode, and that
-cswitches to UTF-8/codepoints mode, and the known issues with this.
And slightly off topic, I think
-ccan be a bit misleading. I had to read the-hline few times to make sure that-cdisables the C locale and enabled unicode :)Maybe add also
-uto enable unicode/UTF-8 (and deprecate-c?) ? The-uflag seems to be unused in both POSIX and gnu awk.And that page also specifies the POSIX locale chars, which, again, are ASCII, or at least have identical values as ASCII for the vast majority of chars.
where do you see specific encoded values required for specific characters? https://pubs.opengroup.org/onlinepubs/9799919799/basedefs/V1_chap06.html does include the UCS value of each character but i'm pretty sure that's just for reference purposes
Interesting. Are there such systems which can be emulated today? I wouldn't mind doing some tests with C code on such system, to test portability (same for systems where byte is not 8 bits).
i'm not sure i know any that are easily available, i think z/os is an example but i'm not super informed on the system (it's exclusively for mainframes and not open source)
And that page also specifies the POSIX locale chars, which, again, are ASCII, or at least have identical values as ASCII for the vast majority of chars.
where do you see specific encoded values required for specific characters? https://pubs.opengroup.org/onlinepubs/9799919799/basedefs/V1_chap06.html does include the UCS value of each character but i'm pretty sure that's just for reference purposes
Indeed I can't find out whether the UCS values are required or not.
This might be another hint, from POSIX locale collation order:
LC_COLLATE # This is the minimum input for the POSIX locale definition for the # LC_COLLATE category. Characters in this list are in the same order # as in the ASCII codeset.But that's collation order. I think it's still technically possible that they have random values but still sort/collate in this order... (e.g. in regexp bracket
[a-z]).Bottom line, indeed I cannot find a direct line of deduction about their values, but a lot of hints all seem to refer specifically to ASCII values/order/etc...
the C/POSIX locale character encoding doesn't have to be ASCII! ...
That being said, re-reading your comment again, does this (POSIX locale is or isn't a superset of ASCII) carry any relevance to this discussion? or to gOawk regardless of this discussion?
As for my suggestion to replace
-cwith-u, leave it out for now. I'll open a new issue to discuss this and related subjects (TL;DR:-ufor unicode,-bfor byte/binary/C (a-la gawk), and try to auto-detect and activate byte/UTF-8 mode from env LC_* values).That being said, re-reading your comment again, does this (POSIX locale is or isn't a superset of ASCII) carry any relevance to this discussion? or to gOawk regardless of this discussion?
it's tangentially relevant but not much more than as a historical curiosity, definitely not relevant to goawk which assumes ascii regardless
it's tangentially relevant but not much more than as a historical curiosity, definitely not relevant to goawk which assumes ascii regardless
Right, so good thing we can put this asside, because it might never end ;)
Anyway, so the bottom line is that we know of the following without
-c, i.e. in bytes/binary mode:- Bug:
/^.$/always matches a single UTF-8 codepoints, and/^...$/never matches a 3-bytes UTF-8 codepoint. This may/likely extends to any single-char matching and not only these specific examples, maybe also with character ranges/classes. - Bug: Precision in printf format
%.Nscounts codepoints rather than bytes. - Correct: printf format
%cpicks a single byte from a string argument - even if it's a valid UTF-8 codepoint.
It may or may not be simple to adress the precision issue.
It would likely be hard to make the regexp work in bytes mode, which is why we keep this issue open for now.
- Bug:
there's also a similar issue with
%s's precision and field width options:$ goawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' ╋:╋ $ gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' ╋:╋ $ LC_ALL=C gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' [garbled output]curiously enough
%cdoesn't have this problem:$ goawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }' [garbled output] $ gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }' ╋:╋ $ LC_ALL=C gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }' [garbled output]Of course you'll get garbled output with
LC_ALL=C gawkwhen you write it as
"%3s:%.1s"%3s is full string, with spaces left-padded to minimum 3 bytes long. %.1s is just left most byte.
echo "\u2326" | gawk -be '{ printf("[%10s]::[%.10s]\n", $1, $1)'[ ⌦]::[⌦]On the other hand, "%.*s" is my goto way to slice open multi-byte UTF-8 characters in gawk Unicode mode. Why would anyone want that ? For starters, URL % encoding or Base64 encoding.
The only encoding-related goawk thing I could find is at the
goawk -hhelp:So my understanding was that it works in the C/POSIX locale by default (logically) regardless of the system locale (
LC_*values), and using-cswitches to some "Unicode mode", where "Unicode mode" probably means UTF-8.Is this understanding correct?
If yes (and also if no), then it's probably worth documenting more explicitly someplace (the README maybe?), and additionally I think I found an issue that matching always behave as if
-cwas specified:(EDIT: for clarity, moved the
printfandgoawkarguments into$FMTand$ACMD, respectively)This seems to confirm that LC_ALL=C is ignored (my default locale is
en_US.UTF-8), as it doesn't affect the result, at least of this test case.-cdoes affect the result, at least oflength($1)for a 3-bytes single-UTF8-codepoint. It's 3 without-c, and1with-c. So far looks OK.However, the match result is unaffected by
-cas far as I can tell.^.$always matches a 3-bytes codepoint regardless if-cis used or not used, and^...$never matches the same 3 bytes.