Repository navigation
Reduce GitHub Actions usage (DataFusion is currently the top ASF consumer of GHA minutes) #25148
Description
Activity
- addeddevelopment-processRelated to development process of DataFusionRelated to development process of DataFusion
on Sep 10, 2026 CI time breakdown by workflow
Breaking down the report screenshot above, DataFusion is as
Total usage: 336,347 minutes, or 33 full-time runners — over the ASF policy ceiling of 250,000 minutes / 25 full-time runners per calendar week.
Three workflows exceed 5% of total time, and together account for ~87% of all usage:
1. CI — 49.31% (~165,900 min)
- Repo:
apache/datafusion-comet - Source: https://github.com/apache/datafusion-comet/blob/main/.github/workflows/ci.yml
- Comet's top-level CI orchestrator
2. Rust — 27.16% (~91,350 min)
- Repo:
apache/datafusion - Source: https://github.com/apache/datafusion/blob/main/.github/workflows/rust.yml
- The main DataFusion core test suite that runs on every PR push, again in the merge queue, and again on push to
main.
3. Datafusion extended tests — 10.85% (~36,500 min)
- Repo:
apache/datafusion - Source: https://github.com/apache/datafusion/blob/main/.github/workflows/extended.yml
- The long-running correctness suite (extended tests, hash-collision tests, sqlite sqllogictests) that runs on every push to
main/branch-*and on every push to any PR touchingdatafusion/physical*,expr*,optimizer, orsql— a large fraction of all PRs.
Everything else (< 2.5% each)
Workflow Share Dev 2.38% Dependencies 2.01% Ballista Rust 1.73% Delta Contrib Build Gate 1.47% Detect breaking changes 0.96% TPC-H SF10 0.88% CodeQL 0.76% TPC-DS SF1 0.47% PyArrow UDF Tests 0.42% Docker 0.29% h2o 0.24% Security audit 0.22% (other builds) 0.86% Observations
- Unlike the May 2026 (#22455) where Comet was ~78% of usage, the core
datafusionrepo now accounts for ~40%+ of usage itself - The "Datafusion extended tests" slice (10.85%) is exactly the one the merge-queue gating idea in Move Extended tests to the standard ones? #21241 would help
- Repo:
FYI @blaginin
Reacted by Dmitrii BlagininAmong other things the fact we have a merge queue and the jobs timeout likely exacerbates the problem (as each PR has to get put in the queue a few times rerunning CI each time)
I also filed a ticket in Comet
Thanks @alamb for tracking this. Comet got lots of contributions last month and that explains
This ticket may also help DF CI #25092, very frequently the PR removed from MQ because of flaky MinIO tests
Is it possible for Apple to chime in for datafusion and datafusion-comet a little?
Are we getting GitHub Actions minutes from GitHub or from Apache? because if this is from Apache, we might be better change to use Hetzner servers (which is what we are doing in my company)Is it possible for Apple to chime in for datafusion and datafusion-comet a little? Are we getting GitHub Actions minutes from GitHub or from Apache? because if this is from Apache, we might be better change to use Hetzner servers (which is what we are doing in my company)
Organizational part is bumpy atm, we already tried that. Minutes come from GH actions that has been run using ASF shared resources.
Reacted by Adrian Garcia BadaraccoIs it possible for Apple to chime in for datafusion and datafusion-comet a little? Are we getting GitHub Actions minutes from GitHub or from Apache? because if this is from Apache, we might be better change to use Hetzner servers (which is what we are doing in my company)
I believe the runners are provided by Github as an in-kind donation to the ASF
We could try and crowdsource additional runner minutes (the way arrow did with ursa computing, etc) but then someone has to manage / pay / coordinate for those resources
I think we shoudl try and reduce runner usage first, and then if we are still having problems figure out a way to get more resources
Could we have someone (e.g. Apple, probably needs to be someone w/ a big AWS account) run
runs-onrunners in their AWS account and we coordinate GitHub App installation / auth w/ ASF Infra? Then the AWS account (and it's bill) would be long to whoever is funding it and there would be no need to coordinate payment (which I understand can be the hard part).Reacted by Dmitrii BlagininAnother option which works for Apache Spark and tested in Comet as well is running CI on the contributor forks apache/datafusion-comet#4570
It changes the model and also contributor machines are 20-30% slower than ASF but the benefit is ASF resource pool is being used much more considerate. This might be later option, but taking into account DF ecosystem is growing and take more resources we might think of it
Reacted by Adrian Garcia Badaracco and Andrew LambI think some of the benefit of runs-on is limited by the fact that a lot of contributors open PRs from forks, so the
varscontext is empty (or coming from the fork repo) and we still use the slow gh-provided runners.so when you open pr from the fork, the base repo still runs them on the head repo infra. it can be runs-on in DF's case - but i had to disable it in a bunch of places because of costs...
Could we have someone (e.g. Apple, probably needs to be someone w/ a big AWS account) run runs-on runners in their AWS account and we coordinate GitHub App installation / auth w/ ASF Infra?
that seems super reasonable
12 remaining items
@alamb sorry, I missed this. Yes, that was the idea: five of the seven
cargo check <crate> featuresjobs had no dependency cache step, so each one started from a cold build. @kumarUjjawal's #25231 has since added the cache to all five, and #25249 follows up by sharing the workspace check's build output.The other part of that comment was the substrait check job, which is still missing the
runs-on/actionstep that the otherruns-onjobs have. That fix is #25195, and it merges cleanly with #25249.@andygrove for comparison, DataFusion is partway there. Since #25203, the three extended test suites are skipped on PR runs and run in the merge queue instead.
rust.ymlhasn't been split yet: the latest PR run and merge queue run both ran 26 jobs. A PR tier and a queue tier like the ones in apache/datafusion-comet#5843 could be the next step for it.Reacted by Andrew LambComet now uses a merge queue and we run fewer workflows on PRs. We are continuing to move more jobs to the merge queue and find other ways to reduce CI usage.
Thank you @andygrove I will see if we can do something similar here.
Reacted by Andrew LambHere is the report (login required) for the last 7 days (looks pretty similar to initially), but hopefully we will see an overall decline over the next few days
Reacted by namanjain24-sudo and Dmitrii Blagininbut hopefully we will see an overall decline over the next few days
Merged a new pr #25249 let's see if this helps
Reacted by Andrew Lamb- added a commit that references this issue
on Sep 15, 2026 Inspired by the work in Comet i opened a new PR #25365
- added 2 commits that reference this issue
on Sep 16, 2026 I pulled the per job timings for
rust.ymlfrom the Actions API to see where the minutes sit after #25203, #25231 and #25249, and the split by runner turned out to be the whole story. Seven days to 2026-09-16, 738 runs with job timings (496pull_request, 133merge_group, 109push), 19,603 jobs:runs GitHub hosted RunsOn pull_request60,697 min 369 min merge_group4,074 min 6,285 min push3,226 min 4,982 min 89% of the GitHub hosted minutes this workflow spends come from PR runs, and they are hosted because
vars.USE_RUNS_ONis empty for a fork. 486 of the 496 PR runs ran entirely on GitHub hosted runners. The 10 that used RunsOn are all branches in this repository:dependabot/cargo/main/dirs-7,dependabot/github_actions/...and so on. Every fork PR I sampled got none:kosiew/datafusion,kumarUjjawal/datafusion,rluvaton/datafusion,lyne7-sc/datafusion,1fanwang/datafusion. That is the fork point @AdamGS and @blaginin raised, measured: about 122 GitHub hosted minutes per PR run, 60,697 per week.For comparison, pushes to
main, which #25365 targets, are 3,226 GitHub hosted minutes a week, and the merge queue is 4,074. Both are worth having, and neither moves the 60,697.Where the PR minutes go, GitHub hosted only, seven days:
job minutes avg cargo test (amd64)7,185 14.7 cargo examples (amd64)6,155 12.6 verify benchmark results (amd64)4,721 9.6 clippy3,633 7.4 cargo check datafusion-substrait features3,476 7.1 cargo test doc (amd64)3,347 6.8 cargo check datafusion features2,891 5.9 linux build test2,798 5.7 cargo test (macos-aarch64)2,680 5.3 One more number: 111 of the 496 PR runs were cancelled, mostly by the concurrency group when a branch is pushed again. Those minutes are already spent when the cancel lands.
So the lever that matters is making PR runs from forks use the same runners as the queue, rather than trimming what runs after a merge. The two approaches discussed above both do that: give the fork's own runners the work, as Comet does, or route fork PRs through this repository's RunsOn account. If it helps, I can put the collection script somewhere so the numbers can be refreshed rather than taken on trust.
- added a commit that references this issue
on Sep 17, 2026 - added a commit that references this issue
on Sep 18, 2026 Although the recent numbers for the last 7 days are withing the ASF policy boundary (limits: 250,000 minutes / calendar week, 216,000 minutes / any 5-day window). I feel if datafusion continue to see new contributors and the PR volumes keep increasing, we will pass the limit again, that's why I think the fork pr job runs are the way to go forward so we don't have to deal with this again.
I have opened a pr for that: #25591
Reacted by Andrew Lamb


Is your feature request related to a problem or challenge?
Recently we have been seeing the Merge Queue get stuck:
We enquired with ASF Infrastructure on INFRA-28374 and they told us the issue is that the ASF's GitHub Action queue is full which is why it is taking a long time to start jobs:
Note the "datafusion" project as measured by ASF Infra includes all DataFusion subprojects (
datafusion,datafusion-comet,datafusion-ballista,datafusion-python, etc.), so reducing usage may require a cross-repository effort.Here is the report of our current usage (login required):
History
This is the second time we have been contacted by ASF Infra about GHA usage.
The first was in May 2026, when ASF Infra identified DataFusion as a top-5 consumer: #22455. After some tweaks, by June 2026, our usage dropped below the policy thresholds and DataFusion fell out of the top 5.
Useful resources
#project-workflow-optimisationsRelated issues