Repository navigation
fix(persist): keep the shared id counter across restore - #1529
Merged
Merged
Conversation
Snapshots now record the idgen counter per provider and restore advances it, so resources created after a restart or snapshot load no longer reuse restored ids. Old snapshots fall back to a capped scan of their ids. Also persist the Azure DNS etag sequence and mint event source mapping ids as UUIDs.
…wind the Azure DNS etag sequence
NitinKumar004
marked this pull request as ready for review
October 10, 2026 14:05
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #1491 (tracker row PERSIST-X1).
Summary
After a persist restore (
serve --persistrestart,POST /_cloudemu/snapshot,cloudemu snapshot load, time-travel rewind, orpersist.RestoreAllin the library), new resources reused the ids of restored ones and overwrote them. This keeps the shared id counter across a restore.Root cause
internal/idgenmints everyGenerateIDid (i-,vpc-,sg-,ANPA, S3 version ids, ...) and every OCID suffix from one process-wide counter. Snapshots never saved it, so a fresh process started counting from 1 again while the restored stores already held those ids, and the next create replaced a restored resource in memstore. Per-mock counters (EC2 vol/snap/ami, Azure VM, GCP compute, ...) were already saved in each mock's snapshot; the shared one was missed.Fix
idgen.Counter()andidgen.AdvanceTo(n).AdvanceTois a CAS loop that only moves the counter forward, so loading an older snapshot never rewinds it.persist.ProviderStategainsidCounter, read after the provider's services are captured.Restoreadvances the counter before restoring;RestoreAll*advances it for every provider before restoring any of them. Every entry point (serve startup, admin snapshot load, CLI, time travel, libraryExport/RestoreandExportAll/RestoreAll) goes through these, so library and serve behave the same.idCounter. For those, restore does a best-effort scan: it takes the last 8 hex characters of every hex run of 8 or more, ignores values at or above 2^28 (random hex), and advances past the highest. Schema version is unchanged; the field is additive.persist.Exportoutput now carries the counter, soserver/aws/authz_unknown_op_test.go(which compares state before and after a call) zeroes it alongside the CloudTrail log.esm-N) that were not saved either. They are now random UUIDs, the shape Lambda returns. Nothing parsed theesm-Nform; the Azure and GCP mappings are library-only.Tests
internal/idgen: AdvanceTo moves forward, never rewinds, and stays unique under concurrent GenerateID (run with-race).persist: round-trip per id family that resets the counter (fresh process), restores, creates again and checks the new id differs and both resources are in state, for both current and legacy snapshots: AWS EC2 instance, VPC, security group, S3 version id, IAM policy, KMS key, Secrets Manager secret, SNS subscription, Route 53 zone, ELBv2 target group, Lambda ESM; Azure VNet, VM, Key Vault secret, Functions ESM; GCP network, GCE instance, Secret Manager secret, Cloud Functions ESM; OCI VCN, identity user, ONS topic. Plus no-rewind, per-providerExport/Restore, and unit tests for the legacy scan. With the fix disabled these fail with "id reused after restore".providers/azure/dns: etag after restore is new and a stale If-Match is rejected (fails without the fix).go build ./...,go vet,go test -raceon idgen, persist, the touched providers, server/serverkit, server/admin, features/timetravel, cmd/cloudemu, server/aws/lambda, server/azure/dns;golangci-lint --new-from-rev=origin/development: 0 issues.E2E evidence
cloudemu servewith all four providers, aws CLI plus Azure ARM, GCP REST and OCI REST via curl.origin/development,
--persistrestart (create, stop, start, create again):This branch, same script:
On origin/development the Azure DNS wire check happened to pass because zone creation writes other record sets that move the sequence; the provider test covers the actual reuse.
Deferrals
blobEventSeqis not saved. It orders blob events and is not a resource key.templateVersioncounters inproviders/azure/virtualmachines/template.goandproviders/gcp/compute/template.goare not saved. They set a launch template's version number, not its key, so a restart can repeat a version number across templates but cannot overwrite one.server/gcp/pubsuback ids,server/azure/cosmosdbetag sequence,server/gcp/{firestore,clouddns,monitoring}operation sequences,server/azure/managedcassandraoperation counter) are short-lived handler state and are not snapshotted.idgen, this covers them.