Skip to content

Protect machine-facing values, and translate in verified batches - #290

Merged
nonprofittechy merged 5 commits into
mainfrom
287-dont-translate-stable-choices
Sep 12, 2026
Merged

nonprofittechy merged 5 commits into
mainfrom
287-dont-translate-stable-choices

Conversation

@nonprofittechy

Copy link
Copy Markdown
Member

Fixes #287. Fixes #285.

Two issues in the translation tool. They landed on one branch because they touch
the same code, but they are independent and the commits are separable.

#287 — machine-facing values must not be translated

Choices and action buttons pair text a user reads with a value the interview
itself consumes. docassemble marks most of those values untranslatable, but not
all of them: a choices: mapping (Other: other), a [value, label] sequence,
the comparison value of show if: {variable: ..., is: ...}, and an action
button's action and arguments all register the machine-facing value as a
translatable segment, so it reaches a translator and gets changed.

This is already shipped damage, not a hypothetical. Regenerating against the
interviews installed on a live server turns up:

file what it says
AssemblyLine/data/sources/translation_es.xlsx plaintiff→demandante, defendant→demandado, started→iniciado, responded→acusado
MAEvictionDefense/.../eviction_es.xlsx True→Cierto, False→Falso, month→mes, summons→Citaciones
HousingCodeChecklist/.../*_es.xlsx ghostwriting→escritura fantasma, attorney→abogado, individual→persona

translation_stable_values.py walks the parsed interview and returns both the
machine-facing values and every string used as display text. translation_file
pins each value that is not also display text: its tr_text is pre-filled with
the original and shaded grey, ahead of the cache lookup so an already-damaged
workbook is repaired on regeneration, and outside the AI batch so no draft is
ever requested for it. Values docassemble never offered for translation get an
explicit row of their own, keeping the pin in the file rather than in this tool.

A value that doubles as a label is deliberately left alone. A bare
choices: [- Red] stores and displays one string and orig_text is the
workbook's key, so pinning it would freeze the visible label without protecting
anything. That case can only be fixed in docassemble.

Verified against four installed interviews: every pinned row comes out matching
its original, no pin collides with display text, and no pinned value is missing
a row.

#285 — one API call per fragment is too slow

The bulk approach that preceded the current code was dropped because rows came
back mixed up, but the code shows that was an implementation bug rather than a
limit of the idea: max_chunk_size was a token budget used as a list index
stride, so chunk 0 got the whole list and every later chunk was an empty slice,
and results.update() then merged whatever came back without looking at it.

The middle ground is small batches whose responses are verified rather than
trusted. plan_batches fills batches greedily in one linear pass, measuring each
fragment once. alignment_problems checks every returned row against its own
source — placeholders, Mako directives, token-length ratio, and whether the text
is verbatim some other fragment's source — and translate_verified re-translates
only the rows that fail, halving the batch until a lone fragment goes back as a
plain-text request, which cannot be misaligned.

Measured against interviews installed on a live server, drafting with
gpt-5.6-luna:

before after
ALAffidavitOfIndigency, 716 rows needing a draft ~716 calls 29 calls, 39s, 0 undrafted, $0.02
40-fragment benchmark 41 calls, 46s 2 calls, 5.4s
150-fragment samples, four interviews, es/zh-t/ht/vi — 6–7 calls each, 8–13s

One first-pass rejection across all four languages, resolved by retry. No row
left misaligned.

Calibrating the verifier

Run over the 11,498 human-translated rows shipped in the packages on that server,
the verifier objects to 10 (0.09%), and on inspection all ten are real defects in
those files — a dropped % if, ${showifdef(...)} reduced to $, and in
eviction_vi.xlsx "My name is misspelled on the summons." translated as
fixed_term, which is the row-mixup this issue describes, already shipped.

That corpus also set two of the rules, both of which started out wrong:

  • Quoted strings inside a placeholder are masked before comparing, because
    ${'do' if claim_jurytrial else 'do not'} correctly becomes
    ${'sí' if claim_jurytrial else 'no'}.
  • Length is compared in tokens rather than characters, because
    "Excessive foot traffic" → "過多客人" is 0.18x the characters but 0.80x the
    tokens.

Cost

Batch size is capped at 25 fragments, a reliability choice rather than a capacity
one, and four requests run at a time. The token budget also respects pricing: the
whole GPT-5.6 and GPT-6 family bills a request at 2x input and 1.5x output once
its input passes 272K tokens, so the planner treats that as a ceiling. In
practice the 128K output cap binds first, since a reply restates every fragment.

The Mako-retry fallback chain drops back through cheaper older models rather
than climbing tiers, and ends at whatever ALToolbox resolves as the provider's
own small model, which is the one entry that means something on a non-OpenAI
endpoint. Blended over a million tokens in and out:

model blended
gpt-5-nano $0.45 in chain
gpt-4.1-nano $0.50 in chain
gpt-5.6-luna $1.40 default
gpt-5.4-nano $1.45 older but dearer, excluded
gpt-4.1 $10.00 older but dearer, excluded
gpt-5.6-terra / sol $14 / $24 excluded

test_translation_fallback_chain fails the build if an entry ever costs more
than the default.

Worth a reviewer's attention

  • The default model changes to gpt-5.6-luna. A server without that
    deployment will get empty drafts rather than an error, which is why the result
    now carries an undrafted_segments count surfaced on the results screen and in
    both API payloads — previously a 404'd model was indistinguishable from a clean
    run.
  • MAX_MAKO_RETRIES goes 3 → 4 so the end of the fallback list is reachable.
    With three entries ahead of it the small model sat at index 3 and a budget of 3
    never got there.
  • is_valid_mako_block now lexes instead of rendering. It runs once per
    drafted fragment, and rendering executed whatever Python the draft happened to
    contain.
  • Two adjacent bugs fixed in passing: the draft-writing loop reassigned row, so
    leftover cache rows overwrote earlier rows whenever AI translation ran
    alongside an existing translation file, and leftover cache rows were never
    added to seen.
  • max_fragments_per_batch and max_parallel_requests are module constants, not
    plumbed through translation_file. Easy to expose if anyone wants them for
    rate-limit tuning.

Testing

85 new unit tests across test_translation_stable_values.py,
test_translation_batching.py and test_translation_fallback_chain.py; 479 pass
in total. mypy and black clean, and scripts/run_da_build_checks.sh exits 0.
Everything above was additionally exercised against a live docassemble server
with the real interviews installed.

🤖 Generated with Claude Code

https://claude.ai/code/session_01BcvXMCXtgaWf5DajChCUz4

nonprofittechy and others added 4 commits September 12, 2026 10:51
Choices and action buttons pair text a user reads with a value the
interview itself consumes: a choice stores its value and code compares
against it, and an action button dispatches its action name and
arguments through action_argument(). docassemble marks most of those
values untranslatable, but not all of them. A `choices:` mapping
(`Other: other`), a `[value, label]` sequence, the comparison value of
`show if: {variable: ..., is: ...}`, and an action button's `action` and
`arguments` all register the machine-facing value as a translatable
segment, so it lands in the workbook and a translator changes it.

This is not hypothetical. Regenerating against the interviews installed
on a live server turns up values that were already translated away in
shipped files -- AssemblyLine's own translation_es.xlsx renders
plaintiff as demandante, defendant as demandado, started as iniciado and
responded as acusado; MAEvictionDefense renders True as Cierto and False
as Falso.

translation_stable_values.py walks the parsed interview and returns both
the machine-facing values and every string used as display text.
translation_file() then pins each value that is not also display text:
its tr_text is pre-filled with the original and shaded grey, ahead of
the cache lookup so an already-damaged workbook is repaired on
regeneration, and outside the GPT batch so no draft is ever requested
for it. Values docassemble never offered for translation get an explicit
row of their own, keeping the pin in the file rather than in this tool.

A value that doubles as a label is deliberately left alone: a bare
`choices: [- Red]` stores and displays one string, and orig_text is the
workbook's key, so pinning it would freeze the visible label without
protecting anything. That one can only be fixed in docassemble.

Two adjacent bugs fixed in passing: the draft-writing loop reassigned
`row`, so leftover cache rows overwrote earlier rows whenever AI
translation ran alongside an existing translation file, and leftover
cache rows were never added to `seen`.

Fixes #287

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BcvXMCXtgaWf5DajChCUz4
A draft used one API call per fragment, so a mid-sized interview took a
couple of thousand round trips. The bulk approach that preceded it was
dropped because rows came back mixed up, but the code shows that was an
implementation bug rather than a limit of the idea: max_chunk_size was a
token budget used as a list index stride, so chunk 0 got the whole list
and every later chunk was an empty slice, and results.update() then
merged whatever came back without looking at it.

The middle ground is small batches whose responses are verified rather
than trusted. translation_batching.plan_batches fills batches greedily in
one linear pass, measuring each fragment once. alignment_problems checks
every returned row against its own source -- placeholders, Mako
directives, token-length ratio, and whether the text is verbatim some
other fragment's source -- and translate_verified re-translates only the
rows that fail, halving the batch until a lone fragment goes back as a
plain-text request, which cannot be misaligned.

Tuned and measured against the interviews installed on a live server,
drafting with gpt-5.6-luna:

  ALAffidavitOfIndigency, 716 rows needing a draft
    before  ~716 calls
    after     29 calls, 39s, 0 undrafted, $0.02

  150-fragment samples, four interviews, es/zh-t/ht/vi
    6-7 calls each, 9-13s, one first-pass rejection across all four,
    resolved by retry; no row left misaligned

Run over the 11,498 human-translated rows shipped in the packages on that
server, the verifier objects to 10 (0.09%), and on inspection all ten are
real defects in those files -- a dropped `% if`, `${showifdef(...)}`
reduced to `$`, and in eviction_vi.xlsx "My name is misspelled on the
summons." translated as "fixed_term". That corpus is also what set two of
the rules: quoted strings inside a placeholder are masked before
comparing, because `${'do' if x else 'do not'}` correctly becomes
`${'sí' if x else 'no'}`, and length is compared in tokens rather than
characters, because "Excessive foot traffic" -> "過多客人" is 0.18x the
characters but 0.80x the tokens.

Batch size is capped at 25 fragments, which is a reliability choice
rather than a capacity one, and four requests run at a time. The token
budget also respects pricing: the whole GPT-5.6 and GPT-6 family bills a
request at 2x input and 1.5x output once its input passes 272K tokens,
and a translation batch has no reason to be anywhere near that, so the
planner treats the threshold as a ceiling. In practice the 128K output
cap binds first, since a reply restates every fragment.

Also here: is_valid_mako_block now lexes instead of rendering. It runs
once per drafted fragment, and rendering executed whatever Python the
draft happened to contain. The default model moves to gpt-5.6-luna with
the rest of its family as the fallback chain, and the result reports an
estimated cost plus a count of rows the AI could not draft, so an
undeployed model no longer looks like a clean run.

Fixes #285

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BcvXMCXtgaWf5DajChCUz4
The Mako-retry chain climbed the 5.6 tiers, so a fragment that kept
coming back with broken template syntax could end up drafted by sol at
17x the cost of the default. What breaks a repeated failure is a
different model rather than a dearer one, so the chain now drops back
through older generations instead.

Blended over a million tokens in and out, which is roughly the shape of a
translation workload:

  gpt-5-nano     $0.45
  gpt-4.1-nano   $0.50
  gpt-5.6-luna   $1.40   (default)
  gpt-5.4-nano   $1.45
  gpt-4.1       $10.00
  gpt-5.6-terra $14.00
  gpt-5.6-sol   $24.00

so the chain is luna -> gpt-4.1-nano -> gpt-5-nano. There is no
gpt-5.5-nano; the nano line runs gpt-5-nano then gpt-5.4-nano. Both
gpt-5.4-nano and gpt-4.1 are older than the default but cost more than
it, so neither is in the chain, and test_translation_fallback_chain now
fails the build if an entry ever costs more than the default.

All four older models are added to MODEL_PRICING, which the batch planner
reads as well as the cost estimate: the 4.1 pair cap output at 32K rather
than 128K, so a batch sized for luna has to be a quarter of that for
them. Prefix matching now prefers the longest key, because
"gpt-4.1-nano-2025-04-14" also starts with "gpt-4.1", which costs 20x as
much per output token.

Checked against the live endpoint: all three chain entries answer, return
every row, and leave `% if` / `% endif` intact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BcvXMCXtgaWf5DajChCUz4
The named chain is three OpenAI models, which is no help to a server
pointed somewhere else: on such a server every entry 404s and the retry
budget runs out having learned nothing. ALToolbox already works out a
"small" model from the docassemble configuration, then the configured
model sets, then the endpoint's own model list, so appending that gives
the chain one entry that means something regardless of provider.

small_model_for_fallback() wraps get_default_model(model_type="small")
and is appended after the named chain, so it is the last thing tried. It
is cached: resolving it calls the models endpoint -- 1.4s on this server,
against 0ms once cached -- and it is consulted once per fragment that
needs a retry.

MAX_MAKO_RETRIES goes from 3 to 4 so the end of the list is actually
reachable. With three entries ahead of it the small model sat at index 3
and a budget of 3 never got there, which would have made the whole thing
decorative. The loop now takes min(MAX_MAKO_RETRIES, len(models_to_try)),
giving each candidate one attempt and no more, and the chain test checks
the arithmetic rather than trusting it.

On the server this was checked against, the small model resolves to
gpt-5-nano, which is already in the named chain, so the list dedupes back
to three.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BcvXMCXtgaWf5DajChCUz4

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Critical and moderate findings remain unresolved in retry verification, directive validation, fallback handling, batching limits, and metrics.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Updates ALDashboard translation generation to preserve machine-facing values and use verified batched AI translation with fallbacks and metrics.

Changes:

  • Protects machine-facing values from translation.
  • Adds batching, validation, retries, fallbacks, and cost metrics.
  • Updates UI/API reporting and adds focused tests.
File summaries
File Summary
docassemble/ALDashboard/translation.py Translation orchestration, preservation, retries, fallbacks, and metrics
docassemble/ALDashboard/translation_stable_values.py Detects machine-facing interview values
docassemble/ALDashboard/translation_batching.py Plans batches, verifies responses, and estimates costs
docassemble/ALDashboard/test/test_translation_stable_values.py Tests value preservation
docassemble/ALDashboard/test/test_translation_fallback_chain.py Tests fallback configuration
docassemble/ALDashboard/test/test_translation_batching.py Tests batching and validation
docassemble/ALDashboard/test/test_interview_linter.py Formatting cleanup
docassemble/ALDashboard/interview_linter.py Formatting cleanup
docassemble/ALDashboard/data/questions/generate_translation.yml Displays translation metrics
docassemble/ALDashboard/data/questions/api_translation_background.yml Adds metrics to background API responses
docassemble/ALDashboard/api_dashboard_utils.py Adds metrics to API responses
Review details

Suppressed comments (6)

docassemble/ALDashboard/data/questions/generate_translation.yml:131

  • preserved_values also includes show-if comparison values and action-button actions/arguments, not only choice values, because those are all collected by stable_values_to_preserve(). On interviews using those constructs this results screen label is inaccurate; describe these as machine-facing values instead.
  Number of rows holding a choice value that is kept in the original language (shaded grey, already filled in): ${ translations[index].preserved_values }

docassemble/ALDashboard/translation.py:1169

  • This counter only counts rows whose final value is blank, but translate_with_retries() falls back to original_text after exhausting the models whenever the source parses as Mako. A missing or undeployed model therefore produces the source-language text and leaves undrafted_segments at zero, so the new result/API signal cannot identify the failure described in the PR. Track whether an AI draft was actually produced, or do not treat the original-text fallback as a draft for this count.
            undrafted_segments = sum(
                1
                for row_number, _original_text, _source_language in hold_for_draft_translation
                if not final_translations.get(row_number)
            )

docassemble/ALDashboard/translation.py:102

  • This backend default is bypassed by the dashboard path: generate_translation.yml still defaults the explicitly passed model to gpt-5-nano, so normal UI runs continue using the old model and never use gpt-5.6-luna as the advertised default. Update the dashboard choice/default (and its label), or scope this default change explicitly to API callers.
DEFAULT_TRANSLATION_MODEL = "gpt-5.6-luna"

docassemble/ALDashboard/translation.py:415

  • A row that failed alignment_problems (for example because its placeholder was dropped) is translated again here, but the singleton result is returned without running the verifier again. A one-row request prevents row-ID mixups, not missing placeholders, directives, or empty output; translation_file only rechecks Mako syntax, so this bad retry can be written to the workbook. Re-run alignment_problems on the singleton result and only merge it when it passes.
        if len(retry) == 1:
            row_number, text = retry[0]
            log(f"Re-translating row {row_number} on its own: {problems[row_number]}")
            try:
                good.update(translate_one(row_number, text))

docassemble/ALDashboard/translation.py:373

  • Duplicate response IDs are silently overwritten in translated. A response containing all requested IDs plus a duplicate, such as 0, 1, 1, produces a complete dictionary and passes alignment, while the last duplicate can replace row 1's translation. Track IDs while parsing and invalidate/retry the batch when one appears more than once.
        for segment in segments:
            if not isinstance(segment, dict):
                continue
            segment_id = str(segment.get("id", ""))
            if segment_id not in wanted:
                continue
            value = segment.get("translation")
            if isinstance(value, str):
                translated[int(segment_id)] = value.rstrip()

docassemble/ALDashboard/translation_batching.py:312

  • Masking every quoted literal makes machine-facing arguments look like user-visible conditional text. A fragment such as ${ action_button_html(url_action("save_changes"), label="Save") } can come back with "save_changes" translated and still match this skeleton, so the verifier accepts a broken action. Restrict masking to literals known to be display text, or preserve code arguments while allowing conditional display strings.
# A quoted string inside a placeholder is user-visible text that a translator is
# right to translate: `${'do' if x else 'do not'}` properly becomes
# `${'sí' if x else 'no'}`. Compare the code around the quotes, not the quotes.
_QUOTED = re.compile(r"'[^']*'|\"[^\"]*\"")
  • Files reviewed: 11/11 changed files
  • Comments generated: 5
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread docassemble/ALDashboard/translation.py Outdated
Comment thread docassemble/ALDashboard/translation_batching.py Outdated
Comment thread docassemble/ALDashboard/translation.py
Comment thread docassemble/ALDashboard/translation_batching.py Outdated
Comment thread docassemble/ALDashboard/translation_batching.py
@nonprofittechy
nonprofittechy requested a lite review from Copilot September 12, 2026 16:34
@nonprofittechy
nonprofittechy merged commit 1986477 into main Sep 12, 2026
7 checks passed
@nonprofittechy
nonprofittechy deleted the 287-dont-translate-stable-choices branch September 12, 2026 16:41

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Unresolved critical and moderate findings remain, including compatibility, batching verification, and fallback/reporting issues.

Get a fresh assessment by requesting another Copilot review.

Review details

Suppressed comments (6)

docassemble/ALDashboard/data/questions/generate_translation.yml:129

  • total_rows now includes the extra pinned machine-value rows, while untranslated_segments deliberately excludes them. The result screen therefore understates the percentage: 10 untranslated rows among 10 translatable rows plus 90 pinned rows is displayed as 10%, not 100%. Use a denominator that excludes preserved_values (or expose a translatable-row count) so the percentage agrees with the displayed untranslated-row count.
  Percentage of rows that are not translated: %${ translations[index].untranslated_segments/translations[index].total_rows * 100 }

docassemble/ALDashboard/data/questions/generate_translation.yml:135

  • undrafted_segments also counts the failure path where all retries fail Mako but the original text is valid and is returned as the fallback (translation.py:1175-1182). Those cells contain the original text rather than being blank, so this status message misreports the generated workbook; describe them as rows without a verified AI draft (or mention the original-text fallback).
  **Rows the AI could not draft (left blank for a human): ${ translations[index].undrafted_segments }**

docassemble/ALDashboard/translation.py:1187

  • When all retries fail Mako validation but the original source is valid, the branch above returns original_text to the workbook. This line is the only place a row is marked as AI-drafted, so that fallback row is counted as undrafted_segments even though its cell is populated with the source text, contradicting the UI text that undrafted rows are left blank. Track source fallbacks separately or adjust the count/message so the result accurately describes what was written.
                    ai_drafted_rows.add(row_number)

docassemble/ALDashboard/translation.py:1186

  • When the configured model is unavailable, translate_fragments_gpt returns no candidate. Ordinary source text parses as valid Mako, so the earlier fallback branch copies the original text into tr_text instead of leaving the draft blank; undrafted_segments then reports it as undrafted while the UI says it was left blank. Restrict the original-text fallback to a non-empty candidate that failed Mako validation, so an unavailable model remains visibly undrafted.
                            f"Unable to create valid Mako translation for row {row_number}; leaving draft empty."
                        )
                        return ""

docassemble/ALDashboard/translation.py:1379

  • values_to_preserve is keyed only by text, so stable_value represents the first occurrence of a value. If that value occurs in both a target-language question and a non-target question whose value was not emitted in question.translations, this condition skips the only explicit row; the non-target occurrence is never written by the main loop and remains unprotected in the workbook. Track occurrences by source language (or skip only when every occurrence is in tr_lang) before suppressing the pinned row.
            if value_text in seen or stable_value.language == tr_lang:

docassemble/ALDashboard/translation_stable_values.py:159

  • This collision check omits the question's other user-visible fields (question, subquestion, under, help, note, and html). If a machine value is also the prompt, such as question: other with a stored choice other, it is not added to display_texts, gets pinned, and the prompt remains in the source language instead of being translated. Collect those parsed display fields as well; the linter's display catalog enumerates the same YAML fields in interview_linter.py:474.
        for attribute in ("content", "subcontent", "helptext"):
            content_text = _text_object_source(getattr(question, attribute, None))
            if content_text:
                display_texts.add(content_text)
  • Files reviewed: 11/11 changed files
  • Comments generated: 5
  • Review effort level: Lite

Comment on lines +500 to +506
preserved_values: int = (
0 # Number of rows holding a machine-facing choice value, pre-filled with the original text
)
undrafted_segments: int = (
0 # Rows AI translation was asked for but could not produce; left blank for a human
)
estimated_cost_usd: float = 0.0 # Rough API cost of the AI drafts, if any
Comment on lines +305 to +307
_MAKO_EXPRESSION = re.compile(r"\$\{.*?\}", re.DOTALL)
_JINJA_TAG = re.compile(r"\{\{.*?\}\}|\{%.*?%\}", re.DOTALL)
_MAKO_DIRECTIVE = re.compile(r"^[ \t]*%(?!%)[^\r\n]*", re.MULTILINE)
Comment on lines +468 to +472
source_tokens = estimate_tokens(source.strip(), model)
if source_tokens >= LENGTH_CHECK_MIN_TOKENS:
ratio = estimate_tokens(stripped, model) / source_tokens
minimum_ratio = 0.15 if _uses_dense_script(stripped) else MIN_LENGTH_RATIO
if ratio < minimum_ratio or ratio > MAX_LENGTH_RATIO:
input type: radio
choices:
- GPT-5 Nano (cheapest, default): gpt-5-nano
- GPT-5.6 Luna (cheapest, default): gpt-5.6-luna
# on the next model in the fallback list. Five accommodates a custom configured
# model, the three named fallbacks, and the provider's own small model. The
# default model needs only four of these attempts.
MAX_MAKO_RETRIES = 5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants