Skip to content

SharePoint Online: speed up ACL sync via concurrent user processing #4435

Description

@MacBlu

Problem Description

The get_access_control() method in the SharePoint Online connector processes users sequentially — each user is fully awaited before the next one begins. For every user who belongs to more than 20 groups (the Microsoft Graph API $expand limit), an additional paginated Graph API call is issued to fetch their full group membership. In enterprise tenants with thousands of users — many of whom are service accounts or admins belonging to dozens of groups — this creates a long chain of serial network round trips, making the ACL sync significantly slower than it needs to be.

Proposed Solution

Replace the sequential async for / await process_user(user) loop in get_access_control() with a bounded concurrent task pool using the existing ConcurrentTasks utility (connectors.utils). Results are collected into an asyncio.Queue and yielded as tasks complete, keeping throughput high without unbounded memory growth.

Key details:

Concurrency bounded at DEFAULT_PARALLEL_CONNECTION_COUNT (10) — matches the existing connection pool size used elsewhere in the connector
The deduplication guard (already_seen_ids) is unchanged — it runs before a task is submitted, so no duplicate ACL documents are produced
Results begin yielding immediately as tasks complete rather than waiting for the full user list to be processed

Alternatives

  1. /users/delta incremental endpoint — Microsoft Graph supports a delta query for users that returns only changed users since a previous sync token. This would reduce subsequent ACL syncs from O(N all users) to O(Δ changed users). Higher impact for repeat syncs but requires cursor management and a schema change to the sync state. Worth pursuing as a follow-up.

  2. $batch API for group membership fallback — For users who exceed the 20-group $expand limit, the current code already issues one extra paged API call per user. These could be batched into a single Graph $batch request (up to 20 sub-requests), similar to how drive_items_permissions_batch() works. This is complementary to concurrent processing and could be a follow-up optimisation.

  3. Increase $top page size — The client already requests $top=999 but the Graph API caps responses at ~100 users per page regardless. No gain available here.

Additional Context

The bottleneck is most severe in enterprise tenants (10 000+ users, many with large group memberships).

ConcurrentTasks is already used for similar fan-out patterns in the Box, Microsoft Teams, Jira, and Confluence connectors — no new infrastructure is introduced.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions