Skip to content

feat(auth): op-level IAM for EKS under --enforce-auth (AUTHZ-X1g-2) - #1540

Draft
NitinKumar004 wants to merge 6 commits into
developmentfrom
feat/authz-eks
Draft

NitinKumar004 wants to merge 6 commits into
developmentfrom
feat/authz-eks

Conversation

@NitinKumar004

Copy link
Copy Markdown
Collaborator

Part of #1495 (AUTHZ-X1g-2, EKS; Route 53 + CloudFront landed in #1536, API Gateway and execute-api:Invoke follow).

What changed

Under cloudemu serve --enforce-auth, EKS is now authorized per operation on resource ARNs instead of at service level, so fine-grained and resource-scoped EKS policies work.

  • Refactor (behaviour-identical). A pure classify(r) (opID, opArgs) that ServeHTTP and IAMChecks / IAMChecksWithContext both use, so the authorized operation and resource are the ones that run. Requests classify cannot name get the handler's existing 4xx with no side effect (fail closed: shortcut principals get the error, policy users 403). 12 new auth-off golden cases were recorded on the base code and pass unchanged.
  • Actions and resources (Service Authorization Reference, Amazon EKS): eks:<Operation>.
    • *: ListClusters, CreateCluster, ListAccessPolicies, DescribeAddonVersions, DescribeAddonConfiguration.
    • cluster ARN: cluster operations, and Create/List of nodegroups, Fargate profiles, add-ons and access entries.
    • nodegroup / Fargate profile / add-on operations: arn:aws:eks:<region>:<account>:<kind>/<cluster>/<name>, built from the names dispatch uses. cloudemu's child ARNs have no trailing id (pre-existing shape), so the checked ARN is exactly the ARN the resource reports, and a missing child is checked on the same shape and still reaches its ResourceNotFoundException (the Terraform destroy waiters rely on this).
    • access entry operations: the entry's stored ARN; a missing entry is checked on access-entry/<cluster>/*.
    • ListUpdates / DescribeUpdate: the nodegroup or add-on named by nodegroupName / addonName, else the cluster.
    • tagging: the resource the ARN names, rebuilt from the names the provider resolves it by.
  • Condition keys. eks:kubernetesVersion, eks:endpointPublicAccess, eks:endpointPrivateAccess, eks:authenticationMode, eks:bootstrapClusterCreatorAdminPermissions, eks:loggingType/<type> (CreateCluster, UpdateClusterConfig, UpdateClusterVersion as the SAR lists them); eks:principalArn, eks:accessEntryType, eks:username, eks:kubernetesGroups (CreateAccessEntry request, and from the stored entry on access-entry operations, plus eks:clusterName, which the SAR lists as an access-entry resource key); eks:policyArn, eks:accessScope, eks:namespaces (Associate / Disassociate); aws:RequestTag/*, aws:TagKeys, aws:ResourceTag/*. The cloudemu-only tags field of UpdateClusterConfig also needs eks:TagResource.
  • Deny body. EKS restJson1 AccessDeniedException through a DenyWriter (X-Amzn-ErrorType plus {"code","message"}, the handler's own error shape).
  • Fixes found while binding the checked ARN to the acted-on resource:
    • F-EKS-1: DescribeUpdate ignored nodegroupName / addonName, so a nodegroup update could be read with cluster permission only. It now finds a nodegroup or add-on update only through that name and a cluster update only without one, as the EKS API Reference describes the parameter (required for a nodegroup or add-on update).
    • Tagging resolved a resource by the names in the ARN alone, so an ARN of another account or region acted on the local resource of the same name, and a bare cluster name was accepted. A tagging ARN must now be an EKS resource ARN of this account and region: another account or region answers NotFoundException, a value that is not one BadRequestException (the error types the EKS tagging API returns).
  • Not in this PR: iam:PassRole for cluster, node, pod execution and service account roles (AUTHZ-X1k). The Kubernetes API of a cluster (/k8s/...) stays authn-only as before.
  • Docs: flag help, EnforceAuth comment, contrib/server README.

Sources

  • Service Authorization Reference, machine-readable: servicereference.us-east-1.amazonaws.com/v1/eks/eks.json (actions, resource types, ARN formats, action and resource condition keys; eks:clusterName is a key of the access-entry resource).
  • AWS What's New, April 2026, "Amazon EKS enhances cluster governance with new IAM condition keys": endpoint access, eks:kubernetesVersion and others on CreateCluster, UpdateClusterConfig and UpdateClusterVersion.
  • EKS API Reference, DescribeUpdate (nodegroupName / addonName required for nodegroup / add-on updates) and TagResource / ListTagsForResource errors (BadRequestException, NotFoundException).

Verification

  • Unit: classify table (every route and error branch), tagging ARN parsing, per-op IAMChecksWithContext table (actions, ARNs, condition keys, missing children), DescribeUpdate resource test, foreign tagging ARN SDK test.
  • Gate: TestAuthzMatrixEKS (scoped cluster and nodegroup ARNs, missing nodegroup in scope gets 404, single Deny, ResourceTag and kubernetesVersion conditions, DescribeUpdate checked on the nodegroup, foreign tagging ARN), TestUnknownRESTOpHasNoSideEffect and TestOpLevelRESTHandlersAreResolvers extended, TestAuthOffResponsesUnchanged (golden recorded on the base code).
  • contrib/server: TestEnforceAuthEKS with the real SDK.
  • Live, cloudemu serve --enforce-auth, scoped IAM user (eks:* on the e2e cluster and its children, Deny UpdateNodegroupVersion):
    • aws CLI: create/describe cluster, nodegroup, add-on, Fargate profile, access entry and policy association, tags OK; DescribeCluster and tags of another cluster denied; UpdateNodegroupVersion denied with the explicit-deny message; DescribeUpdate with --nodegroup-name OK and without it ResourceNotFoundException; DescribeNodegroup of a missing nodegroup ResourceNotFoundException; deletes OK.
    • kubectl via aws eks update-kubeconfig against the cluster endpoint: get and create namespaces work, unaffected.
    • Terraform (hashicorp/aws 6.66.0) as that user with aws_eks_cluster, aws_eks_node_group, aws_eks_addon: apply, plan -detailed-exitcode = 0, in-place nodegroup scaling update (UpdateNodegroupConfig + DescribeUpdate waiter), plan 0 again, destroy with the delete waiters.
    • Auth off: the same Terraform lifecycle and UpdateNodegroupVersion succeed as before.
  • Gates: go build ./..., go vet, go test -race on server/aws/eks, providers/aws/eks, server/aws, server/wire/..., persist; go -C contrib/server test -run Enforce; golangci-lint (new issues) 0; go generate no docs diff.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant