Skip to content

Windows: non-ASCII filenames either percent-escaped or mojibake, depending on --restrict-file-names #388

Description

@zotabee

Hello. I'm using wget2 on Windows and have been unable to get non-ASCII filenames handled correctly with the current release build.

Windows: --restrict-file-names=windows,nocontrol ignores nocontrol; with nocontrol alone, non-ASCII filenames are written as mojibake

Environment

GNU Wget2 2.2.1 - multithreaded metalink/file/website downloader

+digest +https +ssl/gnutls +ipv6 +iri +large-file -nls -ntlm -opie +psl -hsts
-iconv +idn2 +zlib -lzma -brotlidec -zstd -bzip2 -lzip +http2 -gpgme

Windows 11, PowerShell 7.6.4, console codepage 1252.
Installed with winget install --id=GNU.Wget2 -e.

Test case

A server offers a file whose name contains U+00BD (VULGAR FRACTION ONE HALF),
served as Content-Type: text/html; charset=utf-8. The intended name is:

8.<U+00BD>.(1963).DSP.V1.zip

Observed

1. --restrict-file-names=windows,nocontrol -> percent-escaped

wget2 --recursive --level 1 --no-directories --no-parent --accept zip \
      --no-robots --restrict-file-names=windows,nocontrol https://www.sous-titres.eu/films/8_.html
Saving '8.%C2%BD.(1963).DSP.V1.zip'

nocontrol appears to have no effect when combined with windows. The
documentation describes nocontrol as turning off escaping of characters above
127, which is what is being asked for here.

2. --restrict-file-names=nocontrol alone -> mojibake

wget2 --recursive --level 1 --no-directories --no-parent --accept zip \
      --no-robots --restrict-file-names=nocontrol https://www.sous-titres.eu/films/8_.html
Saving '8.<U+00C2><U+00BD>.(1963).DSP.V1.zip'

The escaping is correctly disabled, but the resulting filename on disk is wrong.
<U+00C2><U+00BD> is precisely what the UTF-8 byte sequence C2 BD looks like
when reinterpreted through codepage 1252. So the raw UTF-8 bytes appear to reach
the Windows filesystem API without being converted to UTF-16 first.

NTFS stores filenames as UTF-16 and accepts U+00BD without any problem: the file
can be renamed to the intended form manually and behaves normally afterwards.

3. Same options, name supplied on the command line -> correct

Passing the file URL directly, with the character literal rather than
percent-encoded:

wget2 ... --restrict-file-names=nocontrol "http://<host>/download/<hash>/8.<U+00BD>.(1963).DSP.V1.zip"
wget2 ... --restrict-file-names=nocontrol "https://www.sous-titres.eu/films/download/frahwurvx9m3h0b/8.½.(1963).DSP.V1.zip"
* any download link from the page above
8.<U+00BD>.(1963).DSP.V1.zip        <- correct

and with windows,nocontrol:

8.%BD.(1963).DSP.V1.zip             <- a single byte escaped, not two

This is the same option set as cases 1 and 2, differing only in where the
filename came from. Two things follow:

  • nocontrol clearly does work on its own, which makes its lack of effect in
    windows,nocontrol (case 1) look like a list-parsing or precedence problem
    rather than an unsupported request.
  • %BD is the codepage-1252 single-byte encoding of U+00BD, whereas case 1
    produced %C2%BD, the UTF-8 encoding. So the filename bytes appear to be
    interpreted as being in the local codepage. That assumption holds for a name
    typed on the command line and fails for one taken from a UTF-8 page, which
    would explain the mojibake in case 2.

4. Same page, build WITH iconv (+iconv) -> correct

To isolate whether case 2 is caused by the missing iconv, the same page was
fetched with an MSYS2 UCRT64 build of the same wget2 version:

GNU Wget2 2.2.1 - multithreaded metalink/file/website downloader
+digest +https +ssl/openssl +ipv6 +iri +large-file +nls -ntlm -opie +psl -hsts
+iconv +idn2 +zlib +lzma +brotlidec +zstd +bzip2 -lzip +http2 +gpgme
wget2 --recursive --level 1 --no-directories --no-parent --accept zip \
      --no-robots --restrict-file-names=nocontrol https://www.sous-titres.eu/films/8_.html

Result:

'8.<U+00BD>.(1963).DSP.V1.zip'      <- correct
'8.<U+00BD>.(1963).DSP.V2.zip'
'8.<U+00BD>.(1963).Z2.zip'

Same OS, same wget2 version, same command - only the iconv build flag differs
between this run and case 2. That isolates case 2 to the missing iconv: with
iconv, the filename is converted correctly on save; without it, the raw UTF-8
bytes are written as-is and read back as mojibake.

For reference, a GNU/Linux build (Debian, wget2 2.2.0, +iconv) was also tried
against the same page and likewise produced correct filenames - unsurprising,
since that system's locale is already UTF-8, so remote and local encodings
agree regardless of iconv. The UCRT64 comparison above is the one that isolates
the variable, since Windows' local encoding is not UTF-8.

Expected

8.<U+00BD>.(1963).DSP.V1.zip on disk in case 2, and nocontrol honoured in
case 1.

Confirmed cause, and one remaining question

Case 2 vs case 4 confirms the missing iconv is the cause of the mojibake.
The official Windows release (via winget, GNU.Wget2) is built -iconv. A
same-version, same-OS build with +iconv (MSYS2 UCRT64) fixes it outright.
Building the winget package with iconv enabled would resolve this directly.

Still open: is --restrict-file-names intended to accept a comma-separated
list of modes? If windows,nocontrol (case 1) is valid syntax, nocontrol
appears to have no effect when combined with windows — it works correctly on
its own (cases 2 and 4) but is silently overridden in that combination.

Why it matters

--no-clobber compares filenames, so the chosen spelling has to be stable
across versions. It currently is not: wget 1.99.x escaped these bytes one at a
time in the system codepage (%BD), while 2.x escapes the UTF-8 encoding
(%C2%BD). Refreshing an existing mirror after upgrading therefore downloaded a
second copy of every affected file instead of skipping it, under a second name,
with identical content. In a mirror of ~55,000 files this produced 42 such
duplicate pairs.

Adopting nocontrol to get unescaped names would introduce a third spelling and
the same problem again, and the names it produces are not correct anyway.

Renaming after download is not a workaround for the same reason: the next run no
longer recognises the file and fetches it again.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions