Skip to content

fix(lambda): drop alias-scoped policy, URL and configs on DeleteAlias - #1535

Merged
NitinKumar004 merged 4 commits into
developmentfrom
fix/lambda-alias-delete-cleanup
Oct 11, 2026
Merged

NitinKumar004 merged 4 commits into
developmentfrom
fix/lambda-alias-delete-cleanup

Conversation

@NitinKumar004

@NitinKumar004 NitinKumar004 commented Oct 11, 2026 •

Copy link
Copy Markdown
Collaborator

Part of #1518 (tracker row LAMBDA-X1, auth work in #1495).

Problem

DeleteAlias removed only the alias entry. The state Lambda keys by the alias qualifier stayed behind, so a new alias with the same name picked it all up:

  • the resource-based policy (GetPolicy --qualifier <alias>)
  • the function URL config (GetFunctionUrlConfig --qualifier <alias>)
  • the provisioned concurrency config
  • the event invoke config

Since #1530 the resource policy is used to authorize Invoke under --enforce-auth, so this was a stale grant: a principal granted on the old alias could still invoke the recreated one.

What AWS removes

All four are sub-resources addressed by the alias (or version) qualifier, so they go when the qualifier goes:

Event source mappings are left alone. They are separate resources, and the DeleteFunction docs say to delete them with DeleteEventSourceMapping.

Change

  • providers/aws/lambda/aliases.go: DeleteAlias now holds m.mu and calls a new dropQualifierState, which removes the qualifier's policy, URL config, event invoke config and provisioned concurrency config. The maps are replaced, not edited in place, so a reader holding an older funcData copy is unaffected.
  • providers/aws/lambda/versions.go: DeleteVersion (DeleteFunction with a version Qualifier) already removed the same four maps, but edited them in place. It now uses the shared helper.

The fix is in the provider, so the library and serve behave the same. Persist is unaffected: snapshots read the same maps.

Locking

The cleanup alone was not enough. UpdateFunction (behind UpdateFunctionConfiguration and UpdateFunctionCode) copied the function's state, edited it and wrote it back without m.mu. If DeleteAlias ran in between, the write-back restored the deleted alias's policy and configs, and the stale grant came back. Every write of a function's state now goes through m.mu:

  • UpdateFunction, CreateFunction, DeleteFunction, CreateAlias, UpdateAlias, RegisterHandler and Restore now take it. Versions, policy, URL, concurrency, provisioned concurrency, event invoke config, tags and AWS config already did.
  • Snapshot holds it through the final marshal, because the snapshot shares the policy maps that AddPermission edits in place.
  • TagFunction and UntagFunction copy the tag map before editing it, because GetFunction and ListFunctionTags read it without the lock.
  • None of these call each other while holding the lock, so there is no re-entry.
  • Function-engine calls never run under the lock, because a real engine deploy can take seconds. CreateFunction and code updates reserve the function name under m.mu, call the engine unlocked, then re-lock and apply the change to the entry re-read at that point, never to the copy read before the deploy. DeleteFunction removes the entry and its qualifier state under the lock, then calls the engine Remove with the name still reserved.
  • The reservation logic is in providers/aws/lambda/engine_ops.go. Each reservation has a token, and its release is deferred right after reserving. An engine that panics (recovered by net/http under serve) can no longer leave a name reserved for good. A release only clears its own token.
  • In-progress rules follow the function states doc. While a code update is in flight, UpdateFunctionCode, UpdateFunctionConfiguration, PublishVersion and TagResource get ResourceConflictException. CreateFunction of a reserved name gets the same error.
  • DeleteFunction during a code update is allowed and wins. The update's finalize finds the entry gone, removes the deployment, and returns ResourceNotFoundException instead of bringing the function back.
  • Restore clears all reservations. If a Restore puts an entry under a name whose CreateFunction is still deploying, the create keeps the restored entry, removes its own deployment and returns ResourceConflictException.
  • Policy, URL and other config writes on the function still go through during a deploy, and the committed update keeps them.

Known gap

If the engine Remove fails on DeleteFunction, the function is already gone from the emulator, but the engine deployment may still exist. DeleteFunction returns the Internal error, and nothing retries the Remove. The same applies to the best-effort Remove in the cleanup paths above, which ignores errors.

Tests

  • providers/aws/lambda/qualifier_cleanup_test.go: delete then recreate an alias, and all four sub-resources are NotFound. The unqualified policy survives. Same check for DeleteVersion. TestDeleteAliasNotUndoneByConcurrentUpdate runs 2000 rounds of DeleteAlias racing UpdateFunction behind a start barrier, then recreates the alias and checks it has no policy. Without the lock it fails within a few hundred rounds (round 417 locally, GOMAXPROCS=2).
  • providers/aws/lambda/engine_lock_test.go: an engine whose Deploy and Remove block until released. While each call (create, code update, delete) is held, PolicyStatements, ResolveFunctionURL and GetPolicy on another function return within 100ms, and conflicting changes to the reserved name are refused. A grant added during a code deploy survives the update. Against the previous head the test hangs, because reads wait for the deploy. A separate test races 8 CreateFunction calls for the same name, and exactly one succeeds.
  • providers/aws/lambda/engine_ops_test.go has three tests. In the first, a delete during a code deploy wins: the update returns NotFound, its deployment is removed, and the name can be created again. In the second, an engine whose Deploy or Remove panics leaves no reservation behind, so the next Create, Update, PublishVersion and Delete work; this one fails on the previous head with AlreadyExists. In the third, a Restore during a create keeps the restored entry.
  • server/aws/lambda/alias_delete_cleanup_sdk_test.go (aws-sdk-go-v2): after recreating the alias, GetPolicy and GetFunctionUrlConfig return ResourceNotFoundException. A DeleteFunction Qualifier=1 call also drops that version's policy.
  • server/aws/authz_matrix_lambda_qualifier_test.go: under enforced auth, a user named in the alias policy can invoke it, but after delete and recreate the same user gets AccessDeniedException.

Without the fix, the alias provider test, the SDK test and the authz test all fail.

Live check

cloudemu serve --enforce-auth with the aws CLI: a user with no identity grant for Lambda, named in an add-permission --qualifier prod statement, then delete-alias and create-alias prod:

  • origin/development: get-policy --qualifier prod still returns the old statement, get-function-url-config still returns the old URL, and the user's invoke returns 200.
  • this branch: both lookups return ResourceNotFoundException, and the invoke returns AccessDeniedException.

Terraform (aws_lambda_function + aws_lambda_alias + aws_lambda_permission with qualifier): apply, then plan -detailed-exitcode returns 0, then destroy succeeds.

A recreated alias with the same name inherited the old alias's resource policy, function URL, provisioned concurrency and event invoke config, so under --enforce-auth the old grant still authorized Invoke. DeleteVersion now shares the same cleanup helper.
UpdateFunction, CreateFunction, DeleteFunction, CreateAlias, UpdateAlias, RegisterHandler and Restore wrote funcData back without m.mu, so a concurrent UpdateFunction could restore an alias policy DeleteAlias had just removed. Snapshot now reads under the same lock and tag edits copy the map first.
Create, Update and Delete reserve the function name under m.mu, call the engine unlocked, then apply the result to the entry re-read under m.mu. A create, update or delete of a reserved name gets ResourceConflictException, so reads on other functions no longer wait for a slow deploy.
…ess rules

Release is deferred right after reserving and checked by token, so a panicking engine or a Restore cannot wedge a name. DeleteFunction now wins over an in-flight code update, whose finalize removes the deployment. PublishVersion and TagResource are refused while an update is in progress. A create whose name was restored meanwhile keeps the restored entry.
@NitinKumar004
NitinKumar004 marked this pull request as ready for review October 11, 2026 06:21
@NitinKumar004
NitinKumar004 merged commit e348e61 into development Oct 11, 2026
23 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant