Summary
With more than one backup job in a container, the jobs silence each other's Zabbix monitoring.
dbbackup_post_dbbackup sends a low level discovery payload containing only the job that just ran:
zabbix_sender -c ... -k dbbackup.backup -o '[{"{#NAME}":"'${backup_job_db_host}.${backup_job_db_name}'"}]'
Zabbix treats every LLD receipt as the complete list of discovered entities. Anything not in that list counts as lost. Since each job sends only itself, DB01 reports DB02 as gone, and vice versa.
Steps to reproduce
- Import
zabbix_templates/db_backup.json and link it to a host.
- Send the discovery payload for one job:
[{"{#NAME}":"hostB.db2"}] → 8 items are discovered and enabled.
- Send the discovery payload for a second job, as a second job in the same container would:
[{"{#NAME}":"hostA.db1"}]
What is the expected correct behavior?
All configured backup jobs stay discovered and enabled, regardless of which one reported last.
Relevant logs and/or screenshots
Measured on Zabbix 7.4.14 after step 3, from items / item_discovery (status 0 = enabled, 1 = disabled):
{#NAME}=hostA.db1 items=8 status=0 ts_delete=0
{#NAME}=hostB.db2 items=8 status=1 ts_delete=1788695312 (2026-09-06)
The lost job's items are disabled immediately — the discovery rule's enabled_lifetime_type is 2 — and scheduled for deletion after lifetime=7d. All nine of its triggers go with them:
{#NAME}=hostB.db2 status=1 [hostB.db2] No backups detected in 2 days
{#NAME}=hostB.db2 status=1 [hostB.db2] No backups detected in 3 days
{#NAME}=hostB.db2 status=1 [hostB.db2] No backups detected in 4 days
{#NAME}=hostB.db2 status=1 [hostB.db2] No backups detected in 5 days
{#NAME}=hostB.db2 status=1 [hostB.db2] Backup - Failed with errors
So at any moment only the job that reported last has working alerting; for every other job it is switched off entirely, including the age based triggers. A backup failing in such a job cannot raise an alarm by construction.
Sending both names in a single payload ([{"{#NAME}":"hostA.db1"},{"{#NAME}":"hostB.db2"}]) re-enables both, which confirms the payload contents are the cause.
Environment
- Image version / tag:
nfrastack/db-backup:latest (4.9.2), template from main
- Zabbix: 7.4.14 (server, frontend, MySQL 8.4), reproduced by sending the payloads straight to the trapper port
- Host OS: Docker on Windows 11
Possible fixes
Any fix has to decide what one discovered entity represents, and every option changes {#NAME}, which orphans existing item history — so this probably wants a deliberate choice rather than a quick patch:
- Send the full job list from every job.
dbbackup_bootstrap_variables backup_init NN unsets and rewrites all backup_job_* variables, so a job cannot resolve another job's host and name without destroying its own state; it would need an isolated subshell.
- Send the discovery once at container start, where
dbbackup_create_schedulers already knows every job, and let the per-job runs send item data only. Discovery would then refresh only on restart.
- Template-side mitigation only: set the discovery rule to never disable and never delete lost resources. Keeps the items alive, but genuinely removed jobs then alarm forever.
Related
The same line has a second problem. {#NAME} uses ${backup_job_db_name}, which is set once in dbbackup_bootstrap_variables and never reassigned inside the per-database loops — even though dbbackup_post_dbbackup receives the actual database as $1. With DB_NAME=ALL or a comma separated list, every database therefore writes to the same item set (host.ALL) and overwrites the previous one, so the reported size, timestamp and status belong to whichever database happened to run last.
A comma separated DB_NAME also puts a comma inside the unquoted item key parameter dbbackup.backup.size.[{#NAME}], which I have not verified further.
Both would be fixed together by deriving {#NAME} from the database actually backed up. I am happy to open a PR if you have a preference on which direction to take.
Summary
With more than one backup job in a container, the jobs silence each other's Zabbix monitoring.
dbbackup_post_dbbackupsends a low level discovery payload containing only the job that just ran:Zabbix treats every LLD receipt as the complete list of discovered entities. Anything not in that list counts as lost. Since each job sends only itself,
DB01reportsDB02as gone, and vice versa.Steps to reproduce
zabbix_templates/db_backup.jsonand link it to a host.[{"{#NAME}":"hostB.db2"}]→ 8 items are discovered and enabled.[{"{#NAME}":"hostA.db1"}]What is the expected correct behavior?
All configured backup jobs stay discovered and enabled, regardless of which one reported last.
Relevant logs and/or screenshots
Measured on Zabbix 7.4.14 after step 3, from
items/item_discovery(status0 = enabled, 1 = disabled):The lost job's items are disabled immediately — the discovery rule's
enabled_lifetime_typeis2— and scheduled for deletion afterlifetime=7d. All nine of its triggers go with them:So at any moment only the job that reported last has working alerting; for every other job it is switched off entirely, including the age based triggers. A backup failing in such a job cannot raise an alarm by construction.
Sending both names in a single payload (
[{"{#NAME}":"hostA.db1"},{"{#NAME}":"hostB.db2"}]) re-enables both, which confirms the payload contents are the cause.Environment
nfrastack/db-backup:latest(4.9.2), template frommainPossible fixes
Any fix has to decide what one discovered entity represents, and every option changes
{#NAME}, which orphans existing item history — so this probably wants a deliberate choice rather than a quick patch:dbbackup_bootstrap_variables backup_init NNunsets and rewrites allbackup_job_*variables, so a job cannot resolve another job's host and name without destroying its own state; it would need an isolated subshell.dbbackup_create_schedulersalready knows every job, and let the per-job runs send item data only. Discovery would then refresh only on restart.Related
The same line has a second problem.
{#NAME}uses${backup_job_db_name}, which is set once indbbackup_bootstrap_variablesand never reassigned inside the per-database loops — even thoughdbbackup_post_dbbackupreceives the actual database as$1. WithDB_NAME=ALLor a comma separated list, every database therefore writes to the same item set (host.ALL) and overwrites the previous one, so the reported size, timestamp and status belong to whichever database happened to run last.A comma separated
DB_NAMEalso puts a comma inside the unquoted item key parameterdbbackup.backup.size.[{#NAME}], which I have not verified further.Both would be fixed together by deriving
{#NAME}from the database actually backed up. I am happy to open a PR if you have a preference on which direction to take.