Skip to content

fix(aws-sqs): emit queue-state CloudWatch gauges and resolve SenderId from caller - #1539

Open
Satyam-Trivedi-ZS wants to merge 2 commits into
stackshy:developmentfrom
Satyam-Trivedi-ZS:fix/aws-sqs-parity-1514
Open

Satyam-Trivedi-ZS wants to merge 2 commits into
stackshy:developmentfrom
Satyam-Trivedi-ZS:fix/aws-sqs-parity-1514

Conversation

@Satyam-Trivedi-ZS

Copy link
Copy Markdown
Collaborator

Summary

Part of #1514. This PR closes the AWS SQS CloudWatch gauge gaps and makes SenderId the caller's IAM unique id, which is what real SQS reports.

Item Status
(Medium) Missing CloudWatch metrics: remaining SQS gauges Fixed. Visible, NotVisible, Delayed and AgeOfOldestMessage are now published, plus the FIFO-only ApproximateNumberOfGroupsWithInflightMessages
(Low) ApproximateNumberOfMessagesVisible never emitted Fixed
(Low) ApproximateNumberOfMessagesNotVisible never emitted Fixed. Counts in-flight messages only (received, not yet deleted or expired). Delayed messages are not included
(Low) ApproximateNumberOfMessagesDelayed never emitted Fixed
(Low) ApproximateAgeOfOldestMessage not tracked or emitted Fixed. In seconds. On standard queues, poison-pill messages (received 3 or more times) are left out. The age resets when a message moves to a DLQ
(Low) SenderId always the bare account ID Fixed. Uses the existing awsidentity.Resolver, the same one STS GetCallerIdentity and EKS use
(Low) NumberOfMessagesMoved not emitted on DLQ redrive / StartMessageMoveTask Not a real gap. AWS has no SQS metric with this name (see below), so nothing was added

What real AWS does (sources)

  • Available CloudWatch metrics for Amazon SQS
    • Gives the metric names, the units (Count, and Seconds for the age) and the single QueueName dimension.
    • NotVisible is "in-flight messages that have been received but not yet deleted or expired". Delayed is "delayed and not immediately available".
    • Age rules: standard queues leave out poison pills received three or more times, the age resets on a DLQ move, and FIFO queues don't reorder.
    • Lists the FIFO metric ApproximateNumberOfGroupsWithInflightMessages.
    • The page lists every SQS metric, and NumberOfMessagesMoved is not one of them. For DLQs it recommends ApproximateNumberOfMessagesVisible.
    • Move-task progress is an API field instead: ApproximateNumberOfMessagesMoved in ListMessageMoveTasks and CancelMessageMoveTask, which cloudemu already returns.
  • Monitoring Amazon SQS queues using CloudWatch
    • Metrics are pushed every minute for active queues. A queue is active "if it contains any messages, or if any action accesses it".
    • The metrics page also says "All metrics emit non-negative values only when the queue is active". So the count gauges are published even when they are zero, which lets an alarm on them return to OK.
  • ReceiveMessage – SenderId: "For a user, returns the user ID, for example ABCDEFGHI1JKLMNOPQ23R. For an IAM role, returns the IAM role ID, for example ABCDE1F2GH3I4JK5LMNOP:i-a123b456". This is the same value as the UserId from GetCallerIdentity.

Changes

New file providers/aws/sqs/metrics.go

  • sampleGauges counts, in one pass, the visible, in-flight and delayed messages, the FIFO groups with in-flight messages, and the oldest counted message.
  • emitQueueGauges publishes them on AWS/SQS with the QueueName dimension, in a single PutMetricData call so alarms are evaluated once.
  • The emulator has no once-a-minute publisher, so it samples the gauges after each action that changes a queue's messages:
    • SendMessage, including SNS, EventBridge and S3 deliveries, and after Lambda ESM delivery
    • ReceiveMessage (including empty receives) and ReceiveMessagesWithOptions
    • DeleteMessage, ChangeMessageVisibility and PurgeQueue
    • a redrive into a DLQ, from the receive path or the Lambda ESM path (gauges published for the DLQ)
    • StartMessageMoveTask (gauges published for the source and every destination)
  • Nothing is computed when no monitoring backend is wired.

Metrics are published after the queue lock is released

  • A CloudWatch alarm action can publish to SNS, and SNS can deliver straight back into the same queue. Publishing while holding the lock could therefore deadlock.
  • DeleteMessage used to publish NumberOfMessagesDeleted under the lock. That is fixed too, and TestQueueGaugesPublishOutsideQueueLock covers it.

SenderId

  • driver.SendMessageInput and driver.BatchSendEntry gain an optional SenderID field.
  • The SQS wire handler fills it from the shared awsidentity.Resolver:
    • a long-term access key gives its IAM user id
    • an STS session gives <role id>:<session>
    • any other key gives the same stable synthetic AIDA… id that GetCallerIdentity reports
    • under --enforce-auth, the verified principal is used
  • server/aws/aws.go now builds the resolver at the top of newServer, so STS, EKS and SQS share it.
  • Library callers and service-to-service deliveries carry no caller principal, so they keep the account ID.

No new persisted state. SenderID was already in the SQS snapshot, and the gauges are computed from message state. go generate ./... produces no coverage diff because no operations changed.

Provider Coverage

  • AWS
  • Azure: n/a (these are AWS/SQS CloudWatch metrics and an SQS attribute)
  • GCP: n/a
  • OCI: n/a

Checklist

  • All tests pass (go test -p 2 ./...)
  • Linter passes: the repo-wide run on development already has existing findings. This PR adds no new findings in touched packages: golangci-lint run --new-from-rev=upstream/development over ./providers/aws/sqs/... ./server/aws/sqs/... ./services/messagequeue/... ./server/aws/ reports 0 issues
  • Every provider the change applies to implements the same behavior
  • Integration tests added to cloudemu_test.go: covered by SDK wire tests in server/aws/sqs_metrics_sdk_test.go instead
  • Unit tests added to provider test files
  • Regenerated docs (go generate ./..., no diff)

Test Plan

Provider tests in providers/aws/sqs/metrics_test.go (FakeClock):

  • TestQueueGaugesTrackVisibleInFlightDelayedAndAge: send, delayed send, receive, ChangeVisibility, delete and purge, with exact gauge values and the age in seconds. An empty queue publishes no age.
  • TestQueueGaugesAgeSkipsStandardPoisonPills: the third receive removes the message from the age.
  • TestQueueGaugesFIFOCountsInflightGroups
  • TestQueueGaugesFollowDeadLetterRedriveAndMoveTask: the DLQ gauge rises on redrive, the age resets, and StartMessageMoveTask updates both queues.
  • TestQueueGaugesFollowLambdaESMDeadLetterRedrive
  • TestQueueGaugesPublishOutsideQueueLock: calls back into the queue from inside PutMetricData. The old in-lock publish would deadlock here.
  • TestQueueGaugesNotEmittedWithoutMonitoring
  • TestSenderIDFromCallerElseAccount

SDK wire tests with real aws-sdk-go-v2 clients, in server/aws/sqs_metrics_sdk_test.go:

  • TestSDKSQSQueueGaugesReachCloudWatch: SQS send, delayed send and receive, then CloudWatch ListMetrics and GetMetricStatistics.
  • TestSDKSQSSenderIDMatchesCallerIdentity: SenderId equals STS GetCallerIdentity UserId. After IAM CreateRole and STS AssumeRole, it equals AssumedRoleId (AROA…:worker-1).

Both SDK tests fail on development and pass with this change. The failures on development were:

  • ListMetrics AWS/SQS missing ApproximateNumberOfMessagesVisible (got [NumberOfMessagesReceived NumberOfMessagesSent SentMessageSize])
  • expected "AIDA9F86D081884C7D65" actual "123456789012"

Full checks

  • go build ./... is clean.
  • go test -p 2 ./... passes (exit 0), including the SQS, SNS, Lambda ESM, CloudWatch, STS, EventBridge and S3-notification packages.
  • golangci-lint adds no new findings in the touched packages; the run with --new-from-rev=upstream/development reports 0 issues. Without that flag, the touched SQS packages report 5 findings, all on lines this PR doesn't change: sqs/authz.go goconst and nolintlint, sqs.go:1009 prealloc in collectVisibleMessages, and 2 wsl findings in movetask.go. They are also present on development.

Real user run: cloudemu serve --aws-port 4604 driven with the AWS CLI (--endpoint-url http://127.0.0.1:4604, AWS_ACCESS_KEY_ID=test):

aws sqs create-queue --queue-name cli-dlq
aws sqs create-queue --queue-name cli-src --attributes RedrivePolicy={deadLetterTargetArn:<dlq arn>,maxReceiveCount:1}
aws sqs send-message --queue-url $SRC --message-body hello
aws sqs send-message --queue-url $SRC --message-body later --delay-seconds 600
aws sts get-caller-identity                      # UserId AIDA9F86D081884C7D65
aws sqs receive-message --queue-url $SRC --visibility-timeout 1 --attribute-names SenderId ...
                                                 # SenderId AIDA9F86D081884C7D65 (was 000000000000)
aws cloudwatch list-metrics --namespace AWS/SQS  # now lists ApproximateAgeOfOldestMessage, ...Delayed, ...NotVisible, ...Visible for cli-src
aws cloudwatch get-metric-statistics --namespace AWS/SQS --metric-name <gauge> --dimensions Name=QueueName,Value=cli-src --statistics Maximum Minimum SampleCount ...
   Visible 1/0/3 Count, NotVisible 1/0/3 Count, Delayed 1/0/3 Count, AgeOfOldestMessage max 25 Seconds
sleep 2; aws sqs receive-message --queue-url $SRC     # exceeds maxReceiveCount -> DLQ (DLQ ApproximateNumberOfMessages 1)
aws cloudwatch get-metric-statistics ... Value=cli-dlq  # Visible 1/1/1
aws sqs start-message-move-task --source-arn <dlq arn>  # COMPLETED, ApproximateNumberOfMessagesMoved 1
aws cloudwatch get-metric-statistics ... Value=cli-dlq  # Visible 1/0/2 (back to 0 after the move)
aws cloudwatch list-metrics --namespace AWS/SQS --metric-name NumberOfMessagesMoved   # [] (not an AWS metric)

Notes / known limits

  • Gauges are sampled when something happens on the queue, not on a one-minute timer. If a delay or visibility timeout expires with no further activity, the change shows up at the queue's next action.
  • Not changed here: GetQueueAttributes ApproximateNumberOfMessagesNotVisible still counts delayed messages as not visible, while real AWS counts only in-flight messages. The new gauge uses the AWS definition. The attribute can be fixed separately.
  • The fair-queue metrics (...InQuietGroups, ApproximateNumberOfNoisyGroups) are not emitted, because cloudemu doesn't model noisy groups. The NumberOfDeduplicatedSentMessages counter is not part of this issue's gauge items.
  • Messages delivered by SNS or EventBridge keep the account ID as SenderId. Real AWS reports the delivering service's principal there.

Related Issues

Part of #1514

🤖 Generated with Claude Code

Satyam-Trivedi-ZS and others added 2 commits October 10, 2026 23:30
Publish the AWS/SQS queue-state gauges (ApproximateNumberOfMessagesVisible,
ApproximateNumberOfMessagesNotVisible, ApproximateNumberOfMessagesDelayed,
ApproximateAgeOfOldestMessage and, for FIFO queues,
ApproximateNumberOfGroupsWithInflightMessages) on the QueueName dimension
after every action that changes a queue's messages, including DLQ redrive
and StartMessageMoveTask. Metrics are now published outside the queue lock
so an alarm action that delivers back into the queue cannot deadlock.

SenderId is now the caller's IAM unique id (user id, or role id:session for
an assumed role) resolved by the shared awsidentity.Resolver, matching
GetCallerIdentity's UserId; library and service-to-service sends keep the
account id.

Part of stackshy#1514

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… low

Part of stackshy#1514

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant