Skip to content

matching always behaves as if "-c" was specified #273

Description

@avih

The only encoding-related goawk thing I could find is at the goawk -h help:

Additional GoAWK features:
  -c                use Unicode chars for index, length, match, substr, and %c

So my understanding was that it works in the C/POSIX locale by default (logically) regardless of the system locale (LC_* values), and using -c switches to some "Unicode mode", where "Unicode mode" probably means UTF-8.

Is this understanding correct?

If yes (and also if no), then it's probably worth documenting more explicitly someplace (the README maybe?), and additionally I think I found an issue that matching always behave as if -c was specified:

$ # '\342\225\213' is UTF-8 of U+254B (boxdraw bold horizontal and vertical) 

$ FMT='X\nYYY\n\342\225\213\n'

$ printf "$FMT"
X
YYY
╋

$ ACMD='/^.$/ {print length($1) " /^.$/ " $1}; /^...$/ {print length($1) " /^...$/ " $1}'

$ # --- without -c ---

$ printf "$FMT" | ./goawk "$ACMD"
1 /^.$/ X
3 /^...$/ YYY
3 /^.$/ ╋

$ printf "$FMT" | LC_ALL=C ./goawk "$ACMD"
1 /^.$/ X
3 /^...$/ YYY
3 /^.$/ ╋

$ # --- with -c ---

$ printf "$FMT" | ./goawk -c "$ACMD"
1 /^.$/ X
3 /^...$/ YYY
1 /^.$/ ╋

$ printf "$FMT" | LC_ALL=C ./goawk -c "$ACMD"
1 /^.$/ X
3 /^...$/ YYY
1 /^.$/ ╋

(EDIT: for clarity, moved the printf and goawk arguments into $FMT and $ACMD, respectively)

This seems to confirm that LC_ALL=C is ignored (my default locale is en_US.UTF-8), as it doesn't affect the result, at least of this test case.

-c does affect the result, at least of length($1) for a 3-bytes single-UTF8-codepoint. It's 3 without -c, and 1 with -c. So far looks OK.

However, the match result is unaffected by -c as far as I can tell. ^.$ always matches a 3-bytes codepoint regardless if -c is used or not used, and ^...$ never matches the same 3 bytes.

Activity

  1. triallax commented on Feb 28, 2026

    @triallax
    Contributor

    there's also a similar issue with %s's precision and field width options:

    $ goawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }'
      ╋:╋
    $ gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }'
      ╋:╋
    $ LC_ALL=C gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }'
    [garbled output]
    

    curiously enough %c doesn't have this problem:

    $ goawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }'
    [garbled output]                                                                                                                                                       
    $ gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }'
      ╋:╋
    $ LC_ALL=C gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }'
    [garbled output]
    
  2. triallax commented on Feb 28, 2026

    @triallax
    Contributor

    So my understanding was that it works in the C/POSIX locale by default (logically) regardless of the system locale (LC_* values), and using -c switches to some "Unicode mode", where "Unicode mode" probably means UTF-8.

    your understanding is correct; to use the system locale as strictly required by POSIX, goawk would have to use C wide characters and the relevant libc functions (which it does not, probably for the best!)

    also an extremely pedantic side note

    the C/POSIX locale character encoding doesn't have to be ASCII! the list of characters it's required to support encoding and the other requirements are explained here

    in practice it is ASCII on almost every system nowadays, though i'm aware of a couple systems that still use EBCDIC for the POSIX locale

  3. changed the title [-]matching alway behaves as if "-c" was specified[/-] [+]matching always behaves as if "-c" was specified[/+] on Feb 28, 2026
  4. benhoyt commented on Feb 28, 2026

    @benhoyt
    Owner

    Yes, you're right -- these are inconsistencies in bytes/unicode handling. However, the match one in particular is tricky to fix, because Go's regexp is always UTF-8 (you can't put it in bytes/ASCII mode). I'll leave this issue open to track.

  5. avih commented on Feb 28, 2026

    @avih
    Author

    there's also a similar issue with %s's precision and field width options:

    $ goawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }'
      ╋:╋
    $ gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }'
      ╋:╋
    

    Interesting (took me few seconds to realize that the 2nd line is gnu awk...).

    So this looks correct in gnu awk - assuming the system locale is some UTF-8, but incorrect in gOawk, because it wants to behave in the C locale without -c.

    In Unicode/UTF-8 mode, the %.1s extracts a single codepoint fully, and the %.3s wants to extract 3 codepoints but runs out of string after 1.

    $ LC_ALL=C gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }'
    [garbled output]
    

    For reference, the "garbled output", at least for me, seems correct for the C locale: I get the full 3 bytes of the 1st string, followed by :, followed by the 1st byte of the 2nd string, and then newline:

    $ LC_ALL=C gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }' | od -c
    0000000 342 225 213   : 342  \n
    0000006

    curiously enough %c doesn't have this problem:

    $ goawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }'
    [garbled output]                                                                                                                                                       
    $ gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }'
      ╋:╋
    $ LC_ALL=C gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }'
    [garbled output]
    

    Precision for %c has undefined behavior in both C99 and POSIX 2024, and as far as I can tell it's not augmented in the POSIX awk printf spec. I think in all 3 examples the precision is ignored, which is a valid behavior.

    But it is true that in this case gOawk does behave as intended - %c extracts the first char of the following string, which without -c is expected to be the first byte - which indeed it is.

    And for completion, goawk -c 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }' indeed extracts the full codepoint in both arguments (and the precision is ignored again).

    So bottom line, precision for %s without -c is wrong and still counts in codepoints, but %c is correct and extracts the 1st byte/codepoint depending on -c.

    the C/POSIX locale character encoding doesn't have to be ASCII! the list of characters it's required to support encoding and the other requirements are explained here

    Thanks. I'm familiar with this page. For reference, the same page in POSIX 2024.

    In what sense do you consider it "doesn't have to be ASCII"? the portable character set is ASCII as far as I can tell. Do you mean that like few control codes which I missed are different from ASCII? or that it can be different entirely (and therefore this system is not POSIX)?

    For reference, 7.2 POSIX locale says this:

    Conforming systems shall provide a POSIX locale, also known as the C locale. In POSIX.1 the requirements for the POSIX locale are more extensive than the requirements for the C locale as specified in the ISO C standard. However, in a conforming POSIX implementation, the POSIX locale and the C locale are identical.

    And that page also specifies the POSIX locale chars, which, again, are ASCII, or at least have identical values as ASCII for the vast majority of chars.

    i'm aware of a couple systems that still use EBCDIC for the POSIX locale

    Interesting. Are there such systems which can be emulated today? I wouldn't mind doing some tests with C code on such system, to test portability (same for systems where byte is not 8 bits).

    the match one in particular is tricky to fix, because Go's regexp is always UTF-8 (you can't put it in bytes/ASCII mode). I'll leave this issue open to track.

    Yeah, I can understand that. Still worth documenting it someplace, probably at the README. Both the fact that by default it tries to behave in bytes mode, and that -c switches to UTF-8/codepoints mode, and the known issues with this.


    And slightly off topic, I think -c can be a bit misleading. I had to read the -h line few times to make sure that -c disables the C locale and enabled unicode :)

    Maybe add also -u to enable unicode/UTF-8 (and deprecate -c?) ? The -u flag seems to be unused in both POSIX and gnu awk.

  6. triallax commented on Feb 28, 2026

    @triallax
    Contributor

    And that page also specifies the POSIX locale chars, which, again, are ASCII, or at least have identical values as ASCII for the vast majority of chars.

    where do you see specific encoded values required for specific characters? https://pubs.opengroup.org/onlinepubs/9799919799/basedefs/V1_chap06.html does include the UCS value of each character but i'm pretty sure that's just for reference purposes

    Interesting. Are there such systems which can be emulated today? I wouldn't mind doing some tests with C code on such system, to test portability (same for systems where byte is not 8 bits).

    i'm not sure i know any that are easily available, i think z/os is an example but i'm not super informed on the system (it's exclusively for mainframes and not open source)

  7. avih commented on Feb 28, 2026

    @avih
    Author

    And that page also specifies the POSIX locale chars, which, again, are ASCII, or at least have identical values as ASCII for the vast majority of chars.

    where do you see specific encoded values required for specific characters? https://pubs.opengroup.org/onlinepubs/9799919799/basedefs/V1_chap06.html does include the UCS value of each character but i'm pretty sure that's just for reference purposes

    Indeed I can't find out whether the UCS values are required or not.

    This might be another hint, from POSIX locale collation order:

    LC_COLLATE
    # This is the minimum input for the POSIX locale definition for the
    # LC_COLLATE category. Characters in this list are in the same order
    # as in the ASCII codeset.
    

    But that's collation order. I think it's still technically possible that they have random values but still sort/collate in this order... (e.g. in regexp bracket [a-z]).

    Bottom line, indeed I cannot find a direct line of deduction about their values, but a lot of hints all seem to refer specifically to ASCII values/order/etc...

    the C/POSIX locale character encoding doesn't have to be ASCII! ...

    That being said, re-reading your comment again, does this (POSIX locale is or isn't a superset of ASCII) carry any relevance to this discussion? or to gOawk regardless of this discussion?

    As for my suggestion to replace -c with -u, leave it out for now. I'll open a new issue to discuss this and related subjects (TL;DR: -u for unicode, -b for byte/binary/C (a-la gawk), and try to auto-detect and activate byte/UTF-8 mode from env LC_* values).

  8. triallax commented on Mar 1, 2026

    @triallax
    Contributor

    That being said, re-reading your comment again, does this (POSIX locale is or isn't a superset of ASCII) carry any relevance to this discussion? or to gOawk regardless of this discussion?

    it's tangentially relevant but not much more than as a historical curiosity, definitely not relevant to goawk which assumes ascii regardless

  9. avih commented on Mar 1, 2026

    @avih
    Author

    it's tangentially relevant but not much more than as a historical curiosity, definitely not relevant to goawk which assumes ascii regardless

    Right, so good thing we can put this asside, because it might never end ;)

    Anyway, so the bottom line is that we know of the following without -c, i.e. in bytes/binary mode:

    • Bug: /^.$/ always matches a single UTF-8 codepoints, and /^...$/ never matches a 3-bytes UTF-8 codepoint. This may/likely extends to any single-char matching and not only these specific examples, maybe also with character ranges/classes.
    • Bug: Precision in printf format %.Ns counts codepoints rather than bytes.
    • Correct: printf format %c picks a single byte from a string argument - even if it's a valid UTF-8 codepoint.

    It may or may not be simple to adress the precision issue.

    It would likely be hard to make the regexp work in bytes mode, which is why we keep this issue open for now.

  10. mogando668 commented on Sep 18, 2026

    @mogando668

    there's also a similar issue with %s's precision and field width options:

    $ goawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }'
      ╋:╋
    $ gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }'
      ╋:╋
    $ LC_ALL=C gawk 'BEGIN { printf "%3s:%.1s\n", "╋", "╋" }'
    [garbled output]
    

    curiously enough %c doesn't have this problem:

    $ goawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }'
    [garbled output]                                                                                                                                                       
    $ gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }'
      ╋:╋
    $ LC_ALL=C gawk 'BEGIN { printf "%3c:%.1c\n", "╋", "╋" }'
    [garbled output]
    

    Of course you'll get garbled output with

    LC_ALL=C gawk

    when you write it as

    "%3s:%.1s"

    %3s is full string, with spaces left-padded to minimum 3 bytes long. %.1s is just left most byte.

    echo "\u2326" | gawk -be '{ printf("[%10s]::[%.10s]\n", $1, $1)'

    [ ⌦]::[⌦]

    On the other hand, "%.*s" is my goto way to slice open multi-byte UTF-8 characters in gawk Unicode mode. Why would anyone want that ? For starters, URL % encoding or Base64 encoding.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions