Repository navigation
fix(haskell): stop escaped char literals at the closing quote (#2439) - #2547
Zhuoxi2000 wants to merge 2 commits into
Conversation
…eusData#2439) The vendored grammar lexes an escaped char literal with /'\\[^ ]*'/, whose [^ ]* also crosses newlines and quotes. With only newlines between '\n' and a primed name (g'), the char token runs through g', so f spans g' and g' gets no node. On one line, ['\n','\t'] lexes as a single char token with no parse error. Add regression tests for both forms. Signed-off-by: Edson <zhuoxi2000@gmail.com>
…ta#2439) tree-sitter-haskell lexes an escaped char literal with /'\\[^ ]*'/. [^ ]* also matches newlines and quotes, and the lexer keeps the longest match, so the token ran on to the last ' before the next space. In f = '\n' g' = 2 it ran from '\n through g': f spanned three lines, g' got no node and the valid module reported has_error. On one line ['\n','\t'] lexed as a single char token, silently. Patch the generated char lex states in haskell/parser.c: the escape body (state 28) also stops at whitespace, and after the closing quote states 112/114/116 no longer re-enter the body; 114 takes one more ' only to complete '\''. The token is now '\ [^'\s]* ' '? plus up to two #. '\'', '\\', '\x41', '\SOH', '\^A', '\1234' and '\n'# lex as before. The reporter's suspect, take_char_literal in scanner.c, is not involved: that path handles '\n' correctly. Upstream pin 7fa19f195803 and master both carry the regex, so the patch is recorded in vendored/grammars/MANIFEST.md. The MANIFEST.md and parser.c lines in scripts/vendored-checksums.txt are refreshed with scripts/security-vendored.sh --update. Fixes DeusData#2439 Signed-off-by: Edson <zhuoxi2000@gmail.com>
|
Thanks for opening this — it has been seen, and it is queued. This note is automated, but it is not a brush-off: it exists so you know where your PR stands instead of having to guess from silence. Current review status: working through a backlog. What that means for this PR, concretely:
Things that will genuinely speed it up whenever review does happen:
If this fixes a bug, a reproduction we can run is worth more than a description of the symptom. Thanks for contributing, and sorry in advance for the wait. |
icecold009
left a comment
There was a problem hiding this comment.
AI-assisted review using OpenAI Codex and TypeSafe Jev (jev-1.13.0), on behalf of @icecold009. This is independent community feedback, not a maintainer decision.
I reviewed all four changed files and traced the generated lexer transitions for whitespace/newline stopping, escaped apostrophe handling, and the optional # suffix. I found no actionable defect in the patch. Both added regression cases ran in the hosted extraction suite (410 passed, 0 failed) on this head.
The current required PR / ci-ok check is failing because the Windows CLANG64 shard fails daemon_conflict_log_windows_concurrent_appends_are_not_dropped (tests/test_daemon_version.c:686). That failing test is outside the changed files; I could not establish a link to this Haskell parser patch. Please confirm it reproduces on the base or rerun before merge.
|
Drafted with an AI coding assistant (Claude). Thanks @icecold009 for going through the lexer changes. On the red |
What does this PR do?
Fixes #2439
When a Haskell definition contains an escaped char literal (
'\n') and the next definition has a primed name (g'), the first definition swallows the second. With the repro from #2439,fspans lines 3-5 andg'gets no node. The parse also reportshas_erroreven though the module is valid. The same bug has a one-line form:['\n','\t']lexes as a singlechartoken and reports no error.Cause. The scanner is not the culprit. Patching
take_char_literalinscanner.cchanged nothing, because that path already handles'\n'correctly. The fault is in the grammar's internal lexer.grammar/literal.jsdefineschar: choice(/'[^']'/, /'\\[^ ]*'/). In the escape branch,[^ ]*also matches newlines and quotes, and tree-sitter keeps the longest match, so the token runs to the last'before the next space. In the repro that is the prime ing'. The regex is the same at our pin7fa19f195803and on upstream master (97288e585b0b, 2026-09-30), so there is no upstream fix to pull.Change. This is a hand patch to the generated
charlex states ints_lexininternal/cbm/vendored/grammars/haskell/parser.c. The diff is +2/-9.\t..\r, not only at space.#): no longer loop back into the escape body.'and moves to state 115. That keeps'\''a single token.'\[^'\s]*''?, followed by up to two#as before.The patch is recorded as a row in
vendored/grammars/MANIFEST.mdunder "Local source patches".scripts/security-vendored.sh --updaterefreshes theMANIFEST.mdandparser.clines inscripts/vendored-checksums.txt. No other checksum line changes.Literals that still lex as one token, unchanged:
'\'','\\','\x41','\SOH','\^A','\1234','a','"', MagicHash'\n'#, and TH quotes ('foo,''Bar).One trade-off, also noted in the MANIFEST row: an escaped char followed directly by a stray
', as in'\n'', is taken as one token. That input is not valid Haskell.If you would rather fix this in the grammar, the alternative is to narrow the regex in
literal.js, for example to/'\\('|[^'\s]+)'/, and regenerateparser.cwith tree-sitter 0.25.10. I can rework the PR that way, or send the regex change to tree-sitter-haskell as well.Checklist
git commit -s).scripts/check-dco.sh origin/main..HEADreports OK for both commits.make -f Makefile.cbm testhangs on my machine during ASan shadow-memory init, beforemain()runs. That is a host issue (macOS 26.5, Darwin 25.5, Apple clang 17.0.0) and unrelated to this change. An unsanitized runner was used instead,make -f Makefile.cbm -j4 SANITIZE= BUILD_DIR=build/nosan build/nosan/test-runner, and the sanitized run is left to CI. Results with the fix:extraction: 410 passedpipeline: 318 passedlanguage: 241 passedlang_contract: 41 passednode_creation_probe: 80 passedgrammar_regression: 1 passedgrammar_labels: 2 passedgrammar_imports: 1 passedsecurity: 38 passedscripts/security-vendored.sh: passed (1083 files checked,Vendored integrity check passed)tests/test_vendored_integrity_contract.sh: OKmake -f Makefile.cbm lint-cppcheckpassed with cppcheck 2.20.0, built from the tag CI uses, on the pre-rebase base. This PR changes no file inLINT_SRCS, so it cannot change the cppcheck result onmain.--dry-run --Werrorontests/test_extraction.creports one violation. It already exists onmain(theparseJsonBodytest) and is outside the lines this PR adds.lint-no-suppressandscripts/lint-memory-core.pypass.lint-ci.tests/test_extraction.c, and the first commit adds only these tests:haskell_escaped_char_does_not_swallow_primed_defis the Haskell: an escaped char literal followed by a primed name swallows the next definition (vendored tree-sitter-haskell) #2439 repro. It assertsfends on line 3,g'andhexist, and there is no parse error.haskell_escaped_char_list_lexes_each_literalchecks the one-line form. It asserts that['\n','\t','\'','\\','\x41','\SOH','a']gives seven separatechartokens with no parse error.Test evidence
Before (test commit only, on
main268a9d8):After (this PR):
This PR and #2457 (Haskell infix names) both add tests near the Haskell block in
tests/test_extraction.c, so whichever merges second will need a small rebase. I'll handle that.AI assistance: this change was drafted with an AI coding assistant (Claude) and verified locally with the tests above.