Repository navigation
Protect machine-facing values, and translate in verified batches - #290
Conversation
Choices and action buttons pair text a user reads with a value the
interview itself consumes: a choice stores its value and code compares
against it, and an action button dispatches its action name and
arguments through action_argument(). docassemble marks most of those
values untranslatable, but not all of them. A `choices:` mapping
(`Other: other`), a `[value, label]` sequence, the comparison value of
`show if: {variable: ..., is: ...}`, and an action button's `action` and
`arguments` all register the machine-facing value as a translatable
segment, so it lands in the workbook and a translator changes it.
This is not hypothetical. Regenerating against the interviews installed
on a live server turns up values that were already translated away in
shipped files -- AssemblyLine's own translation_es.xlsx renders
plaintiff as demandante, defendant as demandado, started as iniciado and
responded as acusado; MAEvictionDefense renders True as Cierto and False
as Falso.
translation_stable_values.py walks the parsed interview and returns both
the machine-facing values and every string used as display text.
translation_file() then pins each value that is not also display text:
its tr_text is pre-filled with the original and shaded grey, ahead of
the cache lookup so an already-damaged workbook is repaired on
regeneration, and outside the GPT batch so no draft is ever requested
for it. Values docassemble never offered for translation get an explicit
row of their own, keeping the pin in the file rather than in this tool.
A value that doubles as a label is deliberately left alone: a bare
`choices: [- Red]` stores and displays one string, and orig_text is the
workbook's key, so pinning it would freeze the visible label without
protecting anything. That one can only be fixed in docassemble.
Two adjacent bugs fixed in passing: the draft-writing loop reassigned
`row`, so leftover cache rows overwrote earlier rows whenever AI
translation ran alongside an existing translation file, and leftover
cache rows were never added to `seen`.
Fixes #287
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BcvXMCXtgaWf5DajChCUz4
A draft used one API call per fragment, so a mid-sized interview took a
couple of thousand round trips. The bulk approach that preceded it was
dropped because rows came back mixed up, but the code shows that was an
implementation bug rather than a limit of the idea: max_chunk_size was a
token budget used as a list index stride, so chunk 0 got the whole list
and every later chunk was an empty slice, and results.update() then
merged whatever came back without looking at it.
The middle ground is small batches whose responses are verified rather
than trusted. translation_batching.plan_batches fills batches greedily in
one linear pass, measuring each fragment once. alignment_problems checks
every returned row against its own source -- placeholders, Mako
directives, token-length ratio, and whether the text is verbatim some
other fragment's source -- and translate_verified re-translates only the
rows that fail, halving the batch until a lone fragment goes back as a
plain-text request, which cannot be misaligned.
Tuned and measured against the interviews installed on a live server,
drafting with gpt-5.6-luna:
ALAffidavitOfIndigency, 716 rows needing a draft
before ~716 calls
after 29 calls, 39s, 0 undrafted, $0.02
150-fragment samples, four interviews, es/zh-t/ht/vi
6-7 calls each, 9-13s, one first-pass rejection across all four,
resolved by retry; no row left misaligned
Run over the 11,498 human-translated rows shipped in the packages on that
server, the verifier objects to 10 (0.09%), and on inspection all ten are
real defects in those files -- a dropped `% if`, `${showifdef(...)}`
reduced to `$`, and in eviction_vi.xlsx "My name is misspelled on the
summons." translated as "fixed_term". That corpus is also what set two of
the rules: quoted strings inside a placeholder are masked before
comparing, because `${'do' if x else 'do not'}` correctly becomes
`${'sí' if x else 'no'}`, and length is compared in tokens rather than
characters, because "Excessive foot traffic" -> "過多客人" is 0.18x the
characters but 0.80x the tokens.
Batch size is capped at 25 fragments, which is a reliability choice
rather than a capacity one, and four requests run at a time. The token
budget also respects pricing: the whole GPT-5.6 and GPT-6 family bills a
request at 2x input and 1.5x output once its input passes 272K tokens,
and a translation batch has no reason to be anywhere near that, so the
planner treats the threshold as a ceiling. In practice the 128K output
cap binds first, since a reply restates every fragment.
Also here: is_valid_mako_block now lexes instead of rendering. It runs
once per drafted fragment, and rendering executed whatever Python the
draft happened to contain. The default model moves to gpt-5.6-luna with
the rest of its family as the fallback chain, and the result reports an
estimated cost plus a count of rows the AI could not draft, so an
undeployed model no longer looks like a clean run.
Fixes #285
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BcvXMCXtgaWf5DajChCUz4
The Mako-retry chain climbed the 5.6 tiers, so a fragment that kept coming back with broken template syntax could end up drafted by sol at 17x the cost of the default. What breaks a repeated failure is a different model rather than a dearer one, so the chain now drops back through older generations instead. Blended over a million tokens in and out, which is roughly the shape of a translation workload: gpt-5-nano $0.45 gpt-4.1-nano $0.50 gpt-5.6-luna $1.40 (default) gpt-5.4-nano $1.45 gpt-4.1 $10.00 gpt-5.6-terra $14.00 gpt-5.6-sol $24.00 so the chain is luna -> gpt-4.1-nano -> gpt-5-nano. There is no gpt-5.5-nano; the nano line runs gpt-5-nano then gpt-5.4-nano. Both gpt-5.4-nano and gpt-4.1 are older than the default but cost more than it, so neither is in the chain, and test_translation_fallback_chain now fails the build if an entry ever costs more than the default. All four older models are added to MODEL_PRICING, which the batch planner reads as well as the cost estimate: the 4.1 pair cap output at 32K rather than 128K, so a batch sized for luna has to be a quarter of that for them. Prefix matching now prefers the longest key, because "gpt-4.1-nano-2025-04-14" also starts with "gpt-4.1", which costs 20x as much per output token. Checked against the live endpoint: all three chain entries answer, return every row, and leave `% if` / `% endif` intact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BcvXMCXtgaWf5DajChCUz4
The named chain is three OpenAI models, which is no help to a server pointed somewhere else: on such a server every entry 404s and the retry budget runs out having learned nothing. ALToolbox already works out a "small" model from the docassemble configuration, then the configured model sets, then the endpoint's own model list, so appending that gives the chain one entry that means something regardless of provider. small_model_for_fallback() wraps get_default_model(model_type="small") and is appended after the named chain, so it is the last thing tried. It is cached: resolving it calls the models endpoint -- 1.4s on this server, against 0ms once cached -- and it is consulted once per fragment that needs a retry. MAX_MAKO_RETRIES goes from 3 to 4 so the end of the list is actually reachable. With three entries ahead of it the small model sat at index 3 and a budget of 3 never got there, which would have made the whole thing decorative. The loop now takes min(MAX_MAKO_RETRIES, len(models_to_try)), giving each candidate one attempt and no more, and the chain test checks the arithmetic rather than trusting it. On the server this was checked against, the small model resolves to gpt-5-nano, which is already in the named chain, so the list dedupes back to three. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BcvXMCXtgaWf5DajChCUz4
There was a problem hiding this comment.
🟡 Changes recommended
Critical and moderate findings remain unresolved in retry verification, directive validation, fallback handling, batching limits, and metrics.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
Updates ALDashboard translation generation to preserve machine-facing values and use verified batched AI translation with fallbacks and metrics.
Changes:
- Protects machine-facing values from translation.
- Adds batching, validation, retries, fallbacks, and cost metrics.
- Updates UI/API reporting and adds focused tests.
File summaries
| File | Summary |
|---|---|
docassemble/ALDashboard/translation.py |
Translation orchestration, preservation, retries, fallbacks, and metrics |
docassemble/ALDashboard/translation_stable_values.py |
Detects machine-facing interview values |
docassemble/ALDashboard/translation_batching.py |
Plans batches, verifies responses, and estimates costs |
docassemble/ALDashboard/test/test_translation_stable_values.py |
Tests value preservation |
docassemble/ALDashboard/test/test_translation_fallback_chain.py |
Tests fallback configuration |
docassemble/ALDashboard/test/test_translation_batching.py |
Tests batching and validation |
docassemble/ALDashboard/test/test_interview_linter.py |
Formatting cleanup |
docassemble/ALDashboard/interview_linter.py |
Formatting cleanup |
docassemble/ALDashboard/data/questions/generate_translation.yml |
Displays translation metrics |
docassemble/ALDashboard/data/questions/api_translation_background.yml |
Adds metrics to background API responses |
docassemble/ALDashboard/api_dashboard_utils.py |
Adds metrics to API responses |
Review details
Suppressed comments (6)
docassemble/ALDashboard/data/questions/generate_translation.yml:131
preserved_valuesalso includes show-if comparison values and action-button actions/arguments, not only choice values, because those are all collected bystable_values_to_preserve(). On interviews using those constructs this results screen label is inaccurate; describe these as machine-facing values instead.
Number of rows holding a choice value that is kept in the original language (shaded grey, already filled in): ${ translations[index].preserved_values }
docassemble/ALDashboard/translation.py:1169
- This counter only counts rows whose final value is blank, but
translate_with_retries()falls back tooriginal_textafter exhausting the models whenever the source parses as Mako. A missing or undeployed model therefore produces the source-language text and leavesundrafted_segmentsat zero, so the new result/API signal cannot identify the failure described in the PR. Track whether an AI draft was actually produced, or do not treat the original-text fallback as a draft for this count.
undrafted_segments = sum(
1
for row_number, _original_text, _source_language in hold_for_draft_translation
if not final_translations.get(row_number)
)
docassemble/ALDashboard/translation.py:102
- This backend default is bypassed by the dashboard path:
generate_translation.ymlstill defaults the explicitly passedmodeltogpt-5-nano, so normal UI runs continue using the old model and never usegpt-5.6-lunaas the advertised default. Update the dashboard choice/default (and its label), or scope this default change explicitly to API callers.
DEFAULT_TRANSLATION_MODEL = "gpt-5.6-luna"
docassemble/ALDashboard/translation.py:415
- A row that failed
alignment_problems(for example because its placeholder was dropped) is translated again here, but the singleton result is returned without running the verifier again. A one-row request prevents row-ID mixups, not missing placeholders, directives, or empty output;translation_fileonly rechecks Mako syntax, so this bad retry can be written to the workbook. Re-runalignment_problemson the singleton result and only merge it when it passes.
if len(retry) == 1:
row_number, text = retry[0]
log(f"Re-translating row {row_number} on its own: {problems[row_number]}")
try:
good.update(translate_one(row_number, text))
docassemble/ALDashboard/translation.py:373
- Duplicate response IDs are silently overwritten in
translated. A response containing all requested IDs plus a duplicate, such as0, 1, 1, produces a complete dictionary and passes alignment, while the last duplicate can replace row 1's translation. Track IDs while parsing and invalidate/retry the batch when one appears more than once.
for segment in segments:
if not isinstance(segment, dict):
continue
segment_id = str(segment.get("id", ""))
if segment_id not in wanted:
continue
value = segment.get("translation")
if isinstance(value, str):
translated[int(segment_id)] = value.rstrip()
docassemble/ALDashboard/translation_batching.py:312
- Masking every quoted literal makes machine-facing arguments look like user-visible conditional text. A fragment such as
${ action_button_html(url_action("save_changes"), label="Save") }can come back with"save_changes"translated and still match this skeleton, so the verifier accepts a broken action. Restrict masking to literals known to be display text, or preserve code arguments while allowing conditional display strings.
# A quoted string inside a placeholder is user-visible text that a translator is
# right to translate: `${'do' if x else 'do not'}` properly becomes
# `${'sí' if x else 'no'}`. Compare the code around the quotes, not the quotes.
_QUOTED = re.compile(r"'[^']*'|\"[^\"]*\"")
- Files reviewed: 11/11 changed files
- Comments generated: 5
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
There was a problem hiding this comment.
🟡 Changes recommended
Unresolved critical and moderate findings remain, including compatibility, batching verification, and fallback/reporting issues.
Get a fresh assessment by requesting another Copilot review.
Review details
Suppressed comments (6)
docassemble/ALDashboard/data/questions/generate_translation.yml:129
total_rowsnow includes the extra pinned machine-value rows, whileuntranslated_segmentsdeliberately excludes them. The result screen therefore understates the percentage: 10 untranslated rows among 10 translatable rows plus 90 pinned rows is displayed as 10%, not 100%. Use a denominator that excludespreserved_values(or expose a translatable-row count) so the percentage agrees with the displayed untranslated-row count.
Percentage of rows that are not translated: %${ translations[index].untranslated_segments/translations[index].total_rows * 100 }
docassemble/ALDashboard/data/questions/generate_translation.yml:135
undrafted_segmentsalso counts the failure path where all retries fail Mako but the original text is valid and is returned as the fallback (translation.py:1175-1182). Those cells contain the original text rather than being blank, so this status message misreports the generated workbook; describe them as rows without a verified AI draft (or mention the original-text fallback).
**Rows the AI could not draft (left blank for a human): ${ translations[index].undrafted_segments }**
docassemble/ALDashboard/translation.py:1187
- When all retries fail Mako validation but the original source is valid, the branch above returns
original_textto the workbook. This line is the only place a row is marked as AI-drafted, so that fallback row is counted asundrafted_segmentseven though its cell is populated with the source text, contradicting the UI text that undrafted rows are left blank. Track source fallbacks separately or adjust the count/message so the result accurately describes what was written.
ai_drafted_rows.add(row_number)
docassemble/ALDashboard/translation.py:1186
- When the configured model is unavailable,
translate_fragments_gptreturns no candidate. Ordinary source text parses as valid Mako, so the earlier fallback branch copies the original text intotr_textinstead of leaving the draft blank;undrafted_segmentsthen reports it as undrafted while the UI says it was left blank. Restrict the original-text fallback to a non-empty candidate that failed Mako validation, so an unavailable model remains visibly undrafted.
f"Unable to create valid Mako translation for row {row_number}; leaving draft empty."
)
return ""
docassemble/ALDashboard/translation.py:1379
values_to_preserveis keyed only by text, sostable_valuerepresents the first occurrence of a value. If that value occurs in both a target-language question and a non-target question whose value was not emitted inquestion.translations, this condition skips the only explicit row; the non-target occurrence is never written by the main loop and remains unprotected in the workbook. Track occurrences by source language (or skip only when every occurrence is intr_lang) before suppressing the pinned row.
if value_text in seen or stable_value.language == tr_lang:
docassemble/ALDashboard/translation_stable_values.py:159
- This collision check omits the question's other user-visible fields (
question,subquestion,under,help,note, andhtml). If a machine value is also the prompt, such asquestion: otherwith a stored choiceother, it is not added todisplay_texts, gets pinned, and the prompt remains in the source language instead of being translated. Collect those parsed display fields as well; the linter's display catalog enumerates the same YAML fields ininterview_linter.py:474.
for attribute in ("content", "subcontent", "helptext"):
content_text = _text_object_source(getattr(question, attribute, None))
if content_text:
display_texts.add(content_text)
- Files reviewed: 11/11 changed files
- Comments generated: 5
- Review effort level: Lite
| preserved_values: int = ( | ||
| 0 # Number of rows holding a machine-facing choice value, pre-filled with the original text | ||
| ) | ||
| undrafted_segments: int = ( | ||
| 0 # Rows AI translation was asked for but could not produce; left blank for a human | ||
| ) | ||
| estimated_cost_usd: float = 0.0 # Rough API cost of the AI drafts, if any |
| _MAKO_EXPRESSION = re.compile(r"\$\{.*?\}", re.DOTALL) | ||
| _JINJA_TAG = re.compile(r"\{\{.*?\}\}|\{%.*?%\}", re.DOTALL) | ||
| _MAKO_DIRECTIVE = re.compile(r"^[ \t]*%(?!%)[^\r\n]*", re.MULTILINE) |
| source_tokens = estimate_tokens(source.strip(), model) | ||
| if source_tokens >= LENGTH_CHECK_MIN_TOKENS: | ||
| ratio = estimate_tokens(stripped, model) / source_tokens | ||
| minimum_ratio = 0.15 if _uses_dense_script(stripped) else MIN_LENGTH_RATIO | ||
| if ratio < minimum_ratio or ratio > MAX_LENGTH_RATIO: |
| input type: radio | ||
| choices: | ||
| - GPT-5 Nano (cheapest, default): gpt-5-nano | ||
| - GPT-5.6 Luna (cheapest, default): gpt-5.6-luna |
| # on the next model in the fallback list. Five accommodates a custom configured | ||
| # model, the three named fallbacks, and the provider's own small model. The | ||
| # default model needs only four of these attempts. | ||
| MAX_MAKO_RETRIES = 5 |
Fixes #287. Fixes #285.
Two issues in the translation tool. They landed on one branch because they touch
the same code, but they are independent and the commits are separable.
#287 — machine-facing values must not be translated
Choices and action buttons pair text a user reads with a value the interview
itself consumes. docassemble marks most of those values untranslatable, but not
all of them: a
choices:mapping (Other: other), a[value, label]sequence,the comparison value of
show if: {variable: ..., is: ...}, and an actionbutton's
actionandargumentsall register the machine-facing value as atranslatable segment, so it reaches a translator and gets changed.
This is already shipped damage, not a hypothetical. Regenerating against the
interviews installed on a live server turns up:
AssemblyLine/data/sources/translation_es.xlsxplaintiff→demandante,defendant→demandado,started→iniciado,responded→acusadoMAEvictionDefense/.../eviction_es.xlsxTrue→Cierto,False→Falso,month→mes,summons→CitacionesHousingCodeChecklist/.../*_es.xlsxghostwriting→escritura fantasma,attorney→abogado,individual→personatranslation_stable_values.pywalks the parsed interview and returns both themachine-facing values and every string used as display text.
translation_filepins each value that is not also display text: its
tr_textis pre-filled withthe original and shaded grey, ahead of the cache lookup so an already-damaged
workbook is repaired on regeneration, and outside the AI batch so no draft is
ever requested for it. Values docassemble never offered for translation get an
explicit row of their own, keeping the pin in the file rather than in this tool.
A value that doubles as a label is deliberately left alone. A bare
choices: [- Red]stores and displays one string andorig_textis theworkbook's key, so pinning it would freeze the visible label without protecting
anything. That case can only be fixed in docassemble.
Verified against four installed interviews: every pinned row comes out matching
its original, no pin collides with display text, and no pinned value is missing
a row.
#285 — one API call per fragment is too slow
The bulk approach that preceded the current code was dropped because rows came
back mixed up, but the code shows that was an implementation bug rather than a
limit of the idea:
max_chunk_sizewas a token budget used as a list indexstride, so chunk 0 got the whole list and every later chunk was an empty slice,
and
results.update()then merged whatever came back without looking at it.The middle ground is small batches whose responses are verified rather than
trusted.
plan_batchesfills batches greedily in one linear pass, measuring eachfragment once.
alignment_problemschecks every returned row against its ownsource — placeholders, Mako directives, token-length ratio, and whether the text
is verbatim some other fragment's source — and
translate_verifiedre-translatesonly the rows that fail, halving the batch until a lone fragment goes back as a
plain-text request, which cannot be misaligned.
Measured against interviews installed on a live server, drafting with
gpt-5.6-luna:One first-pass rejection across all four languages, resolved by retry. No row
left misaligned.
Calibrating the verifier
Run over the 11,498 human-translated rows shipped in the packages on that server,
the verifier objects to 10 (0.09%), and on inspection all ten are real defects in
those files — a dropped
% if,${showifdef(...)}reduced to$, and ineviction_vi.xlsx"My name is misspelled on the summons." translated asfixed_term, which is the row-mixup this issue describes, already shipped.That corpus also set two of the rules, both of which started out wrong:
${'do' if claim_jurytrial else 'do not'}correctly becomes${'sí' if claim_jurytrial else 'no'}."Excessive foot traffic" → "過多客人" is 0.18x the characters but 0.80x the
tokens.
Cost
Batch size is capped at 25 fragments, a reliability choice rather than a capacity
one, and four requests run at a time. The token budget also respects pricing: the
whole GPT-5.6 and GPT-6 family bills a request at 2x input and 1.5x output once
its input passes 272K tokens, so the planner treats that as a ceiling. In
practice the 128K output cap binds first, since a reply restates every fragment.
The Mako-retry fallback chain drops back through cheaper older models rather
than climbing tiers, and ends at whatever ALToolbox resolves as the provider's
own small model, which is the one entry that means something on a non-OpenAI
endpoint. Blended over a million tokens in and out:
test_translation_fallback_chainfails the build if an entry ever costs morethan the default.
Worth a reviewer's attention
gpt-5.6-luna. A server without thatdeployment will get empty drafts rather than an error, which is why the result
now carries an
undrafted_segmentscount surfaced on the results screen and inboth API payloads — previously a 404'd model was indistinguishable from a clean
run.
MAX_MAKO_RETRIESgoes 3 → 4 so the end of the fallback list is reachable.With three entries ahead of it the small model sat at index 3 and a budget of 3
never got there.
is_valid_mako_blocknow lexes instead of rendering. It runs once perdrafted fragment, and rendering executed whatever Python the draft happened to
contain.
row, soleftover cache rows overwrote earlier rows whenever AI translation ran
alongside an existing translation file, and leftover cache rows were never
added to
seen.max_fragments_per_batchandmax_parallel_requestsare module constants, notplumbed through
translation_file. Easy to expose if anyone wants them forrate-limit tuning.
Testing
85 new unit tests across
test_translation_stable_values.py,test_translation_batching.pyandtest_translation_fallback_chain.py; 479 passin total. mypy and black clean, and
scripts/run_da_build_checks.shexits 0.Everything above was additionally exercised against a live docassemble server
with the real interviews installed.
🤖 Generated with Claude Code
https://claude.ai/code/session_01BcvXMCXtgaWf5DajChCUz4