Skip to content

middleware: parse quoted ContentCharset parameters correctly - #1205

Open
tianrking wants to merge 1 commit into
go-chi:masterfrom
tianrking:codex/charset-media-parameters
Open

tianrking wants to merge 1 commit into
go-chi:masterfrom
tianrking:codex/charset-media-parameters

Conversation

@tianrking

@tianrking tianrking commented Oct 4, 2026 •

Copy link
Copy Markdown

ContentCharset("UTF-8") currently rejects valid Content-Type: text/plain; charset="UTF-8" requests with 415. Its substring extraction also reads charset= inside another quoted parameter rather than selecting the actual charset parameter. Parse the media parameters with the standard library, retaining case-insensitive charset matching and the existing fallback for absent or non-MIME header values.

Normalize HTTP quoted pairs only within quoted strings before MIME parsing. This is needed because Go's MIME parser deliberately preserves some backslash escapes for legacy browser file paths, whereas HTTP quoted pairs represent the following byte. Escaped quotes and literal backslashes stay encoded for the MIME parser to decode once. Request headers and bodies remain untouched, and the existing bodyless-request bypass stays in place.

Validation on exact source 35fea3e25b3e1e02f154fc5bd4213070b309ad15:

  • Original production plus the same final 22 regressions: 14 assertion failures, 8 controls passing, native exit 1. Original native evidence.
  • Fixed source: 22/22 passing with the race detector on Linux and Windows, for Go 1.24.13, 1.25.14, 1.26.8, and 1.27.1. Cases include real HTTP requests with both known-length and chunked bodies. Final native evidence and artifacts.
  • Official make test passes on all eight combinations: router and middleware suites, 57 and 73 top-level tests respectively, with -race. Full go vet ./... also passes; scoped goimports v0.31.0 reports no changes. Every job verifies actual working-file SHA256 before and after all checks, normalized file blobs against this source, and unchanged HEAD/tracked diff/status.

The standard parser has a measurable cost. On one Linux runner with Go 1.24.13, three samples of the direct parser benchmark measured unquoted headers at 99–100 ns/op (32 B, 1 allocation) before and 337 ns/op (344 B, 3 allocations) after. Quoted/quoted-pair/metadata headers measured 329/398/489 ns/op after, compared with approximately 99/102/82 ns/op before; the original implementation rejects the quoted regression inputs. These are parser microbenchmarks, not HTTP latency measurements. The raw samples, benchmark fixture, source/blob IDs, and SHA256 manifests are attached to the separate benchmark run.

The validation workflows and benchmark fixture are isolated to fork helper branches; this PR changes only the production parser and its regression tests.

AI assistance was used to prepare the implementation and tests. The native test results above were executed and checked against the recorded source and file hashes.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant