Skip to content

fix(persist): keep the shared id counter across restore - #1529

Merged
NitinKumar004 merged 2 commits into
developmentfrom
fix/persist-idgen-restore
Oct 10, 2026
Merged

NitinKumar004 merged 2 commits into
developmentfrom
fix/persist-idgen-restore

Conversation

@NitinKumar004

@NitinKumar004 NitinKumar004 commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

Part of #1491 (tracker row PERSIST-X1).

Summary

After a persist restore (serve --persist restart, POST /_cloudemu/snapshot, cloudemu snapshot load, time-travel rewind, or persist.RestoreAll in the library), new resources reused the ids of restored ones and overwrote them. This keeps the shared id counter across a restore.

Root cause

internal/idgen mints every GenerateID id (i-, vpc-, sg-, ANPA, S3 version ids, ...) and every OCID suffix from one process-wide counter. Snapshots never saved it, so a fresh process started counting from 1 again while the restored stores already held those ids, and the next create replaced a restored resource in memstore. Per-mock counters (EC2 vol/snap/ami, Azure VM, GCP compute, ...) were already saved in each mock's snapshot; the shared one was missed.

Fix

  • idgen.Counter() and idgen.AdvanceTo(n). AdvanceTo is a CAS loop that only moves the counter forward, so loading an older snapshot never rewinds it.
  • persist.ProviderState gains idCounter, read after the provider's services are captured. Restore advances the counter before restoring; RestoreAll* advances it for every provider before restoring any of them. Every entry point (serve startup, admin snapshot load, CLI, time travel, library Export/Restore and ExportAll/RestoreAll) goes through these, so library and serve behave the same.
  • Snapshots written before this change have no idCounter. For those, restore does a best-effort scan: it takes the last 8 hex characters of every hex run of 8 or more, ignores values at or above 2^28 (random hex), and advances past the highest. Schema version is unchanged; the field is additive.
  • Azure DNS now saves its etag sequence, so a record set written after a restore cannot get back an etag it had before and a stale If-Match still fails. Restore keeps the higher of the live and saved values, so an older snapshot cannot move it back.
  • persist.Export output now carries the counter, so server/aws/authz_unknown_op_test.go (which compares state before and after a call) zeroes it alongside the CloudTrail log.
  • Event source mapping ids in Lambda, Azure Functions and Cloud Functions came from separate package-level counters (esm-N) that were not saved either. They are now random UUIDs, the shape Lambda returns. Nothing parsed the esm-N form; the Azure and GCP mappings are library-only.

Tests

  • internal/idgen: AdvanceTo moves forward, never rewinds, and stays unique under concurrent GenerateID (run with -race).
  • persist: round-trip per id family that resets the counter (fresh process), restores, creates again and checks the new id differs and both resources are in state, for both current and legacy snapshots: AWS EC2 instance, VPC, security group, S3 version id, IAM policy, KMS key, Secrets Manager secret, SNS subscription, Route 53 zone, ELBv2 target group, Lambda ESM; Azure VNet, VM, Key Vault secret, Functions ESM; GCP network, GCE instance, Secret Manager secret, Cloud Functions ESM; OCI VCN, identity user, ONS topic. Plus no-rewind, per-provider Export/Restore, and unit tests for the legacy scan. With the fix disabled these fail with "id reused after restore".
  • providers/azure/dns: etag after restore is new and a stale If-Match is rejected (fails without the fix).
  • Scoped gates: go build ./..., go vet, go test -race on idgen, persist, the touched providers, server/serverkit, server/admin, features/timetravel, cmd/cloudemu, server/aws/lambda, server/azure/dns; golangci-lint --new-from-rev=origin/development: 0 issues.

E2E evidence

cloudemu serve with all four providers, aws CLI plus Azure ARM, GCP REST and OCI REST via curl.

origin/development, --persist restart (create, stop, start, create again):

ec2: i-00000011 -> i-00000011              FAIL reused
vpc: vpc-00000015 -> vpc-00000015          FAIL reused
sg: sg-0000001d -> sg-0000001d             FAIL reused (old group now reads as g2)
policy: ANPA00000020 -> ANPA00000020       FAIL reused
s3ver: 00000022 -> 00000022                FAIL reused (two versions with the same id)
azure vnet1 after creating vnet2           FAIL gone

This branch, same script:

-- 1. serve --persist restart
ec2: i-00000011 -> i-0000002f
vpc: vpc-00000015 -> vpc-00000033
sg: sg-0000001d -> sg-0000003b
policy: ANPA00000020 -> ANPA0000003e
s3ver: 00000022 -> 00000040
vcn: ocid1.vcn.oc1.iad.aaaaaaaa0000000000000027 -> ocid1.vcn.oc1.iad.aaaaaaaa0000000000000045
ociuser: ocid1.user.oc1..aaaaaaaa000000000000002c -> ocid1.user.oc1..aaaaaaaa000000000000004a
old security group, instance, Azure VNet, GCP network, OCI VCN, S3 version body: intact
Azure DNS: new etag after restore, stale If-Match -> 412
-- 2. GET /_cloudemu/snapshot, POST it into a fresh process: no reuse against either earlier run
-- 3. legacy state file with idCounter removed: no reuse, old resources intact
== failures: 0 (72 checks)

On origin/development the Azure DNS wire check happened to pass because zone creation writes other record sets that move the sequence; the provider test covers the actual reuse.

Deferrals

  • Azure blob blobEventSeq is not saved. It orders blob events and is not a resource key.
  • The package-level templateVersion counters in providers/azure/virtualmachines/template.go and providers/gcp/compute/template.go are not saved. They set a launch template's version number, not its key, so a restart can repeat a version number across templates but cannot overwrite one.
  • Server handler counters (server/gcp/pubsub ack ids, server/azure/cosmosdb etag sequence, server/gcp/{firestore,clouddns,monitoring} operation sequences, server/azure/managedcassandra operation counter) are short-lived handler state and are not snapshotted.
  • S3 in-progress multipart uploads are not persisted by design, so their ids cannot collide with restored state.
  • OCI ONS subscriptions and Object Storage version ids and retention rules (named in AWS cross-cutting: open real-cloud parity gaps #1491) are not on development yet. If they mint through idgen, this covers them.

Snapshots now record the idgen counter per provider and restore advances it, so resources created after a restart or snapshot load no longer reuse restored ids. Old snapshots fall back to a capped scan of their ids. Also persist the Azure DNS etag sequence and mint event source mapping ids as UUIDs.
@NitinKumar004
NitinKumar004 marked this pull request as ready for review October 10, 2026 14:05
@NitinKumar004
NitinKumar004 merged commit 582dbc2 into development Oct 10, 2026
23 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant