Hello. I'm using wget2 on Windows and have been unable to get non-ASCII filenames handled correctly with the current release build.
Windows: --restrict-file-names=windows,nocontrol ignores nocontrol; with nocontrol alone, non-ASCII filenames are written as mojibake
Environment
GNU Wget2 2.2.1 - multithreaded metalink/file/website downloader
+digest +https +ssl/gnutls +ipv6 +iri +large-file -nls -ntlm -opie +psl -hsts
-iconv +idn2 +zlib -lzma -brotlidec -zstd -bzip2 -lzip +http2 -gpgme
Windows 11, PowerShell 7.6.4, console codepage 1252.
Installed with winget install --id=GNU.Wget2 -e.
Test case
A server offers a file whose name contains U+00BD (VULGAR FRACTION ONE HALF),
served as Content-Type: text/html; charset=utf-8. The intended name is:
8.<U+00BD>.(1963).DSP.V1.zip
Observed
1. --restrict-file-names=windows,nocontrol -> percent-escaped
wget2 --recursive --level 1 --no-directories --no-parent --accept zip \
--no-robots --restrict-file-names=windows,nocontrol https://www.sous-titres.eu/films/8_.html
Saving '8.%C2%BD.(1963).DSP.V1.zip'
nocontrol appears to have no effect when combined with windows. The
documentation describes nocontrol as turning off escaping of characters above
127, which is what is being asked for here.
2. --restrict-file-names=nocontrol alone -> mojibake
wget2 --recursive --level 1 --no-directories --no-parent --accept zip \
--no-robots --restrict-file-names=nocontrol https://www.sous-titres.eu/films/8_.html
Saving '8.<U+00C2><U+00BD>.(1963).DSP.V1.zip'
The escaping is correctly disabled, but the resulting filename on disk is wrong.
<U+00C2><U+00BD> is precisely what the UTF-8 byte sequence C2 BD looks like
when reinterpreted through codepage 1252. So the raw UTF-8 bytes appear to reach
the Windows filesystem API without being converted to UTF-16 first.
NTFS stores filenames as UTF-16 and accepts U+00BD without any problem: the file
can be renamed to the intended form manually and behaves normally afterwards.
3. Same options, name supplied on the command line -> correct
Passing the file URL directly, with the character literal rather than
percent-encoded:
wget2 ... --restrict-file-names=nocontrol "http://<host>/download/<hash>/8.<U+00BD>.(1963).DSP.V1.zip"
wget2 ... --restrict-file-names=nocontrol "https://www.sous-titres.eu/films/download/frahwurvx9m3h0b/8.½.(1963).DSP.V1.zip"
* any download link from the page above
8.<U+00BD>.(1963).DSP.V1.zip <- correct
and with windows,nocontrol:
8.%BD.(1963).DSP.V1.zip <- a single byte escaped, not two
This is the same option set as cases 1 and 2, differing only in where the
filename came from. Two things follow:
nocontrol clearly does work on its own, which makes its lack of effect in
windows,nocontrol (case 1) look like a list-parsing or precedence problem
rather than an unsupported request.
%BD is the codepage-1252 single-byte encoding of U+00BD, whereas case 1
produced %C2%BD, the UTF-8 encoding. So the filename bytes appear to be
interpreted as being in the local codepage. That assumption holds for a name
typed on the command line and fails for one taken from a UTF-8 page, which
would explain the mojibake in case 2.
4. Same page, build WITH iconv (+iconv) -> correct
To isolate whether case 2 is caused by the missing iconv, the same page was
fetched with an MSYS2 UCRT64 build of the same wget2 version:
GNU Wget2 2.2.1 - multithreaded metalink/file/website downloader
+digest +https +ssl/openssl +ipv6 +iri +large-file +nls -ntlm -opie +psl -hsts
+iconv +idn2 +zlib +lzma +brotlidec +zstd +bzip2 -lzip +http2 +gpgme
wget2 --recursive --level 1 --no-directories --no-parent --accept zip \
--no-robots --restrict-file-names=nocontrol https://www.sous-titres.eu/films/8_.html
Result:
'8.<U+00BD>.(1963).DSP.V1.zip' <- correct
'8.<U+00BD>.(1963).DSP.V2.zip'
'8.<U+00BD>.(1963).Z2.zip'
Same OS, same wget2 version, same command - only the iconv build flag differs
between this run and case 2. That isolates case 2 to the missing iconv: with
iconv, the filename is converted correctly on save; without it, the raw UTF-8
bytes are written as-is and read back as mojibake.
For reference, a GNU/Linux build (Debian, wget2 2.2.0, +iconv) was also tried
against the same page and likewise produced correct filenames - unsurprising,
since that system's locale is already UTF-8, so remote and local encodings
agree regardless of iconv. The UCRT64 comparison above is the one that isolates
the variable, since Windows' local encoding is not UTF-8.
Expected
8.<U+00BD>.(1963).DSP.V1.zip on disk in case 2, and nocontrol honoured in
case 1.
Confirmed cause, and one remaining question
Case 2 vs case 4 confirms the missing iconv is the cause of the mojibake.
The official Windows release (via winget, GNU.Wget2) is built -iconv. A
same-version, same-OS build with +iconv (MSYS2 UCRT64) fixes it outright.
Building the winget package with iconv enabled would resolve this directly.
Still open: is --restrict-file-names intended to accept a comma-separated
list of modes? If windows,nocontrol (case 1) is valid syntax, nocontrol
appears to have no effect when combined with windows — it works correctly on
its own (cases 2 and 4) but is silently overridden in that combination.
Why it matters
--no-clobber compares filenames, so the chosen spelling has to be stable
across versions. It currently is not: wget 1.99.x escaped these bytes one at a
time in the system codepage (%BD), while 2.x escapes the UTF-8 encoding
(%C2%BD). Refreshing an existing mirror after upgrading therefore downloaded a
second copy of every affected file instead of skipping it, under a second name,
with identical content. In a mirror of ~55,000 files this produced 42 such
duplicate pairs.
Adopting nocontrol to get unescaped names would introduce a third spelling and
the same problem again, and the names it produces are not correct anyway.
Renaming after download is not a workaround for the same reason: the next run no
longer recognises the file and fetches it again.
Hello. I'm using wget2 on Windows and have been unable to get non-ASCII filenames handled correctly with the current release build.
Windows:
--restrict-file-names=windows,nocontrolignoresnocontrol; withnocontrolalone, non-ASCII filenames are written as mojibakeEnvironment
Windows 11, PowerShell 7.6.4, console codepage 1252.
Installed with
winget install --id=GNU.Wget2 -e.Test case
A server offers a file whose name contains U+00BD (VULGAR FRACTION ONE HALF),
served as
Content-Type: text/html; charset=utf-8. The intended name is:Observed
1.
--restrict-file-names=windows,nocontrol-> percent-escapednocontrolappears to have no effect when combined withwindows. Thedocumentation describes
nocontrolas turning off escaping of characters above127, which is what is being asked for here.
2.
--restrict-file-names=nocontrolalone -> mojibakeThe escaping is correctly disabled, but the resulting filename on disk is wrong.
<U+00C2><U+00BD>is precisely what the UTF-8 byte sequenceC2 BDlooks likewhen reinterpreted through codepage 1252. So the raw UTF-8 bytes appear to reach
the Windows filesystem API without being converted to UTF-16 first.
NTFS stores filenames as UTF-16 and accepts U+00BD without any problem: the file
can be renamed to the intended form manually and behaves normally afterwards.
3. Same options, name supplied on the command line -> correct
Passing the file URL directly, with the character literal rather than
percent-encoded:
and with
windows,nocontrol:This is the same option set as cases 1 and 2, differing only in where the
filename came from. Two things follow:
nocontrolclearly does work on its own, which makes its lack of effect inwindows,nocontrol(case 1) look like a list-parsing or precedence problemrather than an unsupported request.
%BDis the codepage-1252 single-byte encoding of U+00BD, whereas case 1produced
%C2%BD, the UTF-8 encoding. So the filename bytes appear to beinterpreted as being in the local codepage. That assumption holds for a name
typed on the command line and fails for one taken from a UTF-8 page, which
would explain the mojibake in case 2.
4. Same page, build WITH iconv (
+iconv) -> correctTo isolate whether case 2 is caused by the missing iconv, the same page was
fetched with an MSYS2 UCRT64 build of the same wget2 version:
Result:
Same OS, same wget2 version, same command - only the
iconvbuild flag differsbetween this run and case 2. That isolates case 2 to the missing iconv: with
iconv, the filename is converted correctly on save; without it, the raw UTF-8
bytes are written as-is and read back as mojibake.
For reference, a GNU/Linux build (Debian, wget2 2.2.0,
+iconv) was also triedagainst the same page and likewise produced correct filenames - unsurprising,
since that system's locale is already UTF-8, so remote and local encodings
agree regardless of iconv. The UCRT64 comparison above is the one that isolates
the variable, since Windows' local encoding is not UTF-8.
Expected
8.<U+00BD>.(1963).DSP.V1.zipon disk in case 2, andnocontrolhonoured incase 1.
Confirmed cause, and one remaining question
Case 2 vs case 4 confirms the missing
iconvis the cause of the mojibake.The official Windows release (via winget,
GNU.Wget2) is built-iconv. Asame-version, same-OS build with
+iconv(MSYS2 UCRT64) fixes it outright.Building the winget package with iconv enabled would resolve this directly.
Still open: is
--restrict-file-namesintended to accept a comma-separatedlist of modes? If
windows,nocontrol(case 1) is valid syntax,nocontrolappears to have no effect when combined with
windows— it works correctly onits own (cases 2 and 4) but is silently overridden in that combination.
Why it matters
--no-clobbercompares filenames, so the chosen spelling has to be stableacross versions. It currently is not: wget 1.99.x escaped these bytes one at a
time in the system codepage (
%BD), while 2.x escapes the UTF-8 encoding(
%C2%BD). Refreshing an existing mirror after upgrading therefore downloaded asecond copy of every affected file instead of skipping it, under a second name,
with identical content. In a mirror of ~55,000 files this produced 42 such
duplicate pairs.
Adopting
nocontrolto get unescaped names would introduce a third spelling andthe same problem again, and the names it produces are not correct anyway.
Renaming after download is not a workaround for the same reason: the next run no
longer recognises the file and fetches it again.