Skip to content

Reduce GitHub Actions usage (DataFusion is currently the top ASF consumer of GHA minutes) #25148

Description

@alamb

Is your feature request related to a problem or challenge?

Recently we have been seeing the Merge Queue get stuck:

We enquired with ASF Infrastructure on INFRA-28374 and they told us the issue is that the ASF's GitHub Action queue is full which is why it is taking a long time to start jobs:

We are aware of a long term usage pattern whereby our limited number of
concurrent GitHub Actions jobs running on GitHub hosted runners is at 100%
utilisation most of the day most of the time. Thus, from time to time we
communicate to projects who utilise the most minutes to review their
workflows for efficiency improvements.

At the moment, datafusion is the top consumer of GitHub Action minutes over
the last 7 days, and 3x the 10th in the list of minutes used.

Note the "datafusion" project as measured by ASF Infra includes all DataFusion subprojects (datafusion, datafusion-comet, datafusion-ballista, datafusion-python, etc.), so reducing usage may require a cross-repository effort.

Here is the report of our current usage (login required):

Image

History

This is the second time we have been contacted by ASF Infra about GHA usage.

The first was in May 2026, when ASF Infra identified DataFusion as a top-5 consumer: #22455. After some tweaks, by June 2026, our usage dropped below the policy thresholds and DataFusion fell out of the top 5.

Useful resources

Related issues

Activity

  1. alamb commented on Sep 10, 2026

    @alamb
    ContributorAuthor

    CI time breakdown by workflow

    Breaking down the report screenshot above, DataFusion is as

    Total usage: 336,347 minutes, or 33 full-time runners — over the ASF policy ceiling of 250,000 minutes / 25 full-time runners per calendar week.

    Three workflows exceed 5% of total time, and together account for ~87% of all usage:

    1. CI — 49.31% (~165,900 min)

    2. Rust — 27.16% (~91,350 min)

    3. Datafusion extended tests — 10.85% (~36,500 min)

    Everything else (< 2.5% each)

    Workflow Share
    Dev 2.38%
    Dependencies 2.01%
    Ballista Rust 1.73%
    Delta Contrib Build Gate 1.47%
    Detect breaking changes 0.96%
    TPC-H SF10 0.88%
    CodeQL 0.76%
    TPC-DS SF1 0.47%
    PyArrow UDF Tests 0.42%
    Docker 0.29%
    h2o 0.24%
    Security audit 0.22%
    (other builds) 0.86%

    Observations

    • Unlike the May 2026 (#22455) where Comet was ~78% of usage, the core datafusion repo now accounts for ~40%+ of usage itself
    • The "Datafusion extended tests" slice (10.85%) is exactly the one the merge-queue gating idea in Move Extended tests to the standard ones? #21241 would help
  2. alamb commented on Sep 10, 2026

    @alamb
    ContributorAuthor
  3. alamb commented on Sep 10, 2026

    @alamb
    ContributorAuthor

    Among other things the fact we have a merge queue and the jobs timeout likely exacerbates the problem (as each PR has to get put in the queue a few times rerunning CI each time)

  4. alamb commented on Sep 10, 2026

    @alamb
    ContributorAuthor
  5. comphead commented on Sep 10, 2026

    @comphead
    Contributor

    Thanks @alamb for tracking this. Comet got lots of contributions last month and that explains

  6. comphead commented on Sep 10, 2026

    @comphead
    Contributor

    This ticket may also help DF CI #25092, very frequently the PR removed from MQ because of flaky MinIO tests

  7. rluvaton commented on Sep 10, 2026

    @rluvaton
    Member

    Is it possible for Apple to chime in for datafusion and datafusion-comet a little?
    Are we getting GitHub Actions minutes from GitHub or from Apache? because if this is from Apache, we might be better change to use Hetzner servers (which is what we are doing in my company)

  8. comphead commented on Sep 10, 2026

    @comphead
    Contributor

    Is it possible for Apple to chime in for datafusion and datafusion-comet a little? Are we getting GitHub Actions minutes from GitHub or from Apache? because if this is from Apache, we might be better change to use Hetzner servers (which is what we are doing in my company)

    Organizational part is bumpy atm, we already tried that. Minutes come from GH actions that has been run using ASF shared resources.

  9. alamb commented on Sep 10, 2026

    @alamb
    ContributorAuthor

    Is it possible for Apple to chime in for datafusion and datafusion-comet a little? Are we getting GitHub Actions minutes from GitHub or from Apache? because if this is from Apache, we might be better change to use Hetzner servers (which is what we are doing in my company)

    I believe the runners are provided by Github as an in-kind donation to the ASF

    We could try and crowdsource additional runner minutes (the way arrow did with ursa computing, etc) but then someone has to manage / pay / coordinate for those resources

    I think we shoudl try and reduce runner usage first, and then if we are still having problems figure out a way to get more resources

  10. adriangb commented on Sep 10, 2026

    @adriangb
    Contributor

    Could we have someone (e.g. Apple, probably needs to be someone w/ a big AWS account) run runs-on runners in their AWS account and we coordinate GitHub App installation / auth w/ ASF Infra? Then the AWS account (and it's bill) would be long to whoever is funding it and there would be no need to coordinate payment (which I understand can be the hard part).

  11. comphead commented on Sep 10, 2026

    @comphead
    Contributor

    Another option which works for Apache Spark and tested in Comet as well is running CI on the contributor forks apache/datafusion-comet#4570

    It changes the model and also contributor machines are 20-30% slower than ASF but the benefit is ASF resource pool is being used much more considerate. This might be later option, but taking into account DF ecosystem is growing and take more resources we might think of it

  12. AdamGS commented on Sep 11, 2026

    @AdamGS
    Contributor

    I think some of the benefit of runs-on is limited by the fact that a lot of contributors open PRs from forks, so the vars context is empty (or coming from the fork repo) and we still use the slow gh-provided runners.

  13. blaginin commented on Sep 11, 2026

    @blaginin
    Member

    so when you open pr from the fork, the base repo still runs them on the head repo infra. it can be runs-on in DF's case - but i had to disable it in a bunch of places because of costs...

  14. blaginin commented on Sep 11, 2026

    @blaginin
    Member

    Could we have someone (e.g. Apple, probably needs to be someone w/ a big AWS account) run runs-on runners in their AWS account and we coordinate GitHub App installation / auth w/ ASF Infra?

    that seems super reasonable

  15. 12 remaining items

  16. namanjain24-sudo commented on Sep 15, 2026

    @namanjain24-sudo
    Contributor

    @alamb sorry, I missed this. Yes, that was the idea: five of the seven cargo check <crate> features jobs had no dependency cache step, so each one started from a cold build. @kumarUjjawal's #25231 has since added the cache to all five, and #25249 follows up by sharing the workspace check's build output.

    The other part of that comment was the substrait check job, which is still missing the runs-on/action step that the other runs-on jobs have. That fix is #25195, and it merges cleanly with #25249.

    @andygrove for comparison, DataFusion is partway there. Since #25203, the three extended test suites are skipped on PR runs and run in the merge queue instead. rust.yml hasn't been split yet: the latest PR run and merge queue run both ran 26 jobs. A PR tier and a queue tier like the ones in apache/datafusion-comet#5843 could be the next step for it.

  17. kumarUjjawal commented on Sep 15, 2026

    @kumarUjjawal
    Contributor

    Comet now uses a merge queue and we run fewer workflows on PRs. We are continuing to move more jobs to the merge queue and find other ways to reduce CI usage.

    Thank you @andygrove I will see if we can do something similar here.

  18. alamb commented on Sep 15, 2026

    @alamb
    ContributorAuthor

    Here is the report (login required) for the last 7 days (looks pretty similar to initially), but hopefully we will see an overall decline over the next few days

    Image

    Here is the chart / details for the past 24 hours
    Image

  19. kumarUjjawal commented on Sep 15, 2026

    @kumarUjjawal
    Contributor

    but hopefully we will see an overall decline over the next few days

    Merged a new pr #25249 let's see if this helps

  20. kumarUjjawal commented on Sep 16, 2026

    @kumarUjjawal
    Contributor

    Inspired by the work in Comet i opened a new PR #25365

  21. namanjain24-sudo commented on Sep 16, 2026

    @namanjain24-sudo
    Contributor

    I pulled the per job timings for rust.yml from the Actions API to see where the minutes sit after #25203, #25231 and #25249, and the split by runner turned out to be the whole story. Seven days to 2026-09-16, 738 runs with job timings (496 pull_request, 133 merge_group, 109 push), 19,603 jobs:

    runs GitHub hosted RunsOn
    pull_request 60,697 min 369 min
    merge_group 4,074 min 6,285 min
    push 3,226 min 4,982 min

    89% of the GitHub hosted minutes this workflow spends come from PR runs, and they are hosted because vars.USE_RUNS_ON is empty for a fork. 486 of the 496 PR runs ran entirely on GitHub hosted runners. The 10 that used RunsOn are all branches in this repository: dependabot/cargo/main/dirs-7, dependabot/github_actions/... and so on. Every fork PR I sampled got none: kosiew/datafusion, kumarUjjawal/datafusion, rluvaton/datafusion, lyne7-sc/datafusion, 1fanwang/datafusion. That is the fork point @AdamGS and @blaginin raised, measured: about 122 GitHub hosted minutes per PR run, 60,697 per week.

    For comparison, pushes to main, which #25365 targets, are 3,226 GitHub hosted minutes a week, and the merge queue is 4,074. Both are worth having, and neither moves the 60,697.

    Where the PR minutes go, GitHub hosted only, seven days:

    job minutes avg
    cargo test (amd64) 7,185 14.7
    cargo examples (amd64) 6,155 12.6
    verify benchmark results (amd64) 4,721 9.6
    clippy 3,633 7.4
    cargo check datafusion-substrait features 3,476 7.1
    cargo test doc (amd64) 3,347 6.8
    cargo check datafusion features 2,891 5.9
    linux build test 2,798 5.7
    cargo test (macos-aarch64) 2,680 5.3

    One more number: 111 of the 496 PR runs were cancelled, mostly by the concurrency group when a branch is pushed again. Those minutes are already spent when the cancel lands.

    So the lever that matters is making PR runs from forks use the same runners as the queue, rather than trimming what runs after a merge. The two approaches discussed above both do that: give the fork's own runners the work, as Comet does, or route fork PRs through this repository's RunsOn account. If it helps, I can put the collection script somewhere so the numbers can be refreshed rather than taken on trust.

  22. kumarUjjawal commented on Sep 22, 2026

    @kumarUjjawal
    Contributor

    Last 7 days report

    Image
  23. kumarUjjawal commented on Sep 22, 2026

    @kumarUjjawal
    Contributor

    Although the recent numbers for the last 7 days are withing the ASF policy boundary (limits: 250,000 minutes / calendar week, 216,000 minutes / any 5-day window). I feel if datafusion continue to see new contributors and the PR volumes keep increasing, we will pass the limit again, that's why I think the fork pr job runs are the way to go forward so we don't have to deal with this again.

    I have opened a pr for that: #25591

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    development-processRelated to development process of DataFusion

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions