Skip to content

lexer: encode \u escapes using locale encoding (rebase) - #1

Closed
Franklin-Qi wants to merge 20 commits into
mainfrom
cursor/pr51-locale-encoding-rebase-cfaf
Closed

Franklin-Qi wants to merge 20 commits into
mainfrom
cursor/pr51-locale-encoding-rebase-cfaf

Conversation

@Franklin-Qi

Copy link
Copy Markdown
Owner

Rebases uutils/awk#51 onto current main and addresses review feedback.

Changes

Review feedback addressed

  • from_locale_name is now a LocaleEncoding constructor method
  • ASCII locale detection uses readable is_ascii closure
  • encode_unicode_escape writes into caller-provided buffer via Extend<u8>
  • char::try_from replaces unwrap() for validation
  • var_os() used for locale environment variables

Testing

  • cargo test -p lexer — 45 tests passed
  • cargo test — all workspace tests passed

Closes: uutils#40

Open in Web Open in Cursor 

Franklin-Qi and others added 20 commits June 17, 2026 12:01
Expand test coverage for the parser and lexer using existing patterns: test_parser! macro assertions for AST s-expressions and token-slice comparisons for the lexer.

Closes: uutils#43
* interpreter: widen instructions, immediates

* lexer, parser: add tokens and nodes for built-in functions

* chore: fix CI

---------

Co-authored-by: yor1xd <gabrielhenriquenf1@gmail.com>
Validate namespace and literal parts of ns::name against gawk's reserved
keyword list when building identifiers, including @namespace directives.

Closes: uutils#37
- Track an anchor at the start of each Pratt subexpression and report errors from that anchor through the offending token, instead of highlighting only the final token.

- Apply this to assignment/index/place errors, prefix operators, unclosed parentheses and brackets, and related Pratt paths. Keep narrow spans for non-associative operators (e.g. chained = ).

- Add unit tests that assert the highlighted source snippets for common parse failures.

Closes: uutils#42

Co-authored-by: Sylvestre Ledru <sylvestre@debian.org>
…orced-non-keyword-names-on-namespaced-identifiers

parser: reject reserved keywords in qualified identifiers
The Pratt parser now handles subtraction vs. unary negative.
Detect charset from LC_ALL/LC_CTYPE/LANG and encode \u sequences into
the locale multibyte encoding (UTF-8, ISO-8859-1, ASCII-only for C/POSIX),
matching gawk. Unrepresentable or invalid code points become '?'.

Closes: uutils#40
@cursor
cursor Bot force-pushed the cursor/pr51-locale-encoding-rebase-cfaf branch from be7c222 to 4105912 Compare June 30, 2026 12:59
@cursor
cursor Bot deleted the cursor/pr51-locale-encoding-rebase-cfaf branch June 30, 2026 13:07
@Alonely0

Alonely0 commented Jun 30, 2026 •

Copy link
Copy Markdown

Would you fix the Git history so the PR only includes your commits? I think the bot screwed up. Otherwise reviewing this is quite hard since the diff is gigantic.

EDIT: nvm I thought this was the main repo, got pinged somehow.

@Franklin-Qi

Franklin-Qi commented Jul 1, 2026 •

Copy link
Copy Markdown
Owner Author

你能修改一下 Git 历史记录,让 PR 只包含你的提交吗?我觉得是机器人搞错了。否则,由于差异巨大,审查起来会非常困难。

编辑:算了,我以为这是主仓库,不知怎么的被通知了。

@Alonely0 Thank you for your feedback. Regarding my issue, the current log is a bit messy, I'm considering closing this PR for now.

@Franklin-Qi Franklin-Qi closed this Jul 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Lexer: Numeric escaping \u, for different locales

3 participants