Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions docs/reference/airflow_tags.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,10 +11,11 @@ More information and the discussions can be found the the original Airflow Tags

### impact/tier tag

We borrow the [tiering system](https://wiki.mozilla.org/Sheriffing/Job_Visibility_Policy#Overview_of_the_Job_Visibility_Tiers) used by our integration and testing sheriffs. This is to maintain a level of consistency across different systems to ensure common language and understanding across teams. Valid tier tags include:
We adapt the [tiering system](https://wiki.mozilla.org/Sheriffing/Job_Visibility_Policy#Overview_of_the_Job_Visibility_Tiers) used by our integration and testing sheriffs. This is to maintain a level of consistency across different systems to ensure common language and understanding across teams. Valid tier tags include:

- **impact/tier_1**: Highest priority/impact/critical DAG. A job with this tag implies that many downstream processes are impacted and affects Mozilla’s (many users across different teams and departments) ability to make decisions. A bug ticket must be created and the issue needs to be resolved as soon as possible.
- **impact/tier_2**: Job of increased importance and impact, however, not critical and only limited impact on other processes. One team or group of people is affected and the pipeline does not generate any business critical metrics. A bug ticket must be created and should be addressed within a few working days.
- **impact/tier_0**: Foundational pipelines that a significant portion of downstream data processing depends on (e.g. copy dedupe, Glean usage). A Tier 0 failure takes priority over all other work — all hands on deck, must fix ASAP. A bug ticket must be created immediately and the issue worked until it is resolved or explicitly handed off to someone who has accepted ownership.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion: The unchanged preamble on line 14 says "We borrow the [tiering system] used by our integration and testing sheriffs... to maintain a level of consistency across different systems to ensure common language and understanding across teams." The linked Job Visibility Policy defines only tiers 1–3, and with this PR the tier_1/tier_2 definitions are also rewritten in terms of data pipelines and SLAs rather than the sheriff definitions. Reword line 14 to say the tiering is inspired by / adapted from the sheriffs' system with a data-pipeline-specific tier 0, so the doc doesn't claim a consistency that no longer holds.

- **impact/tier_1**: Business-critical pipelines: those supporting key daily reporting, important product functionality, or time-sensitive processes such as data sent to external systems or marketing ad platforms. A bug ticket must be created and the issue is expected to be resolved within hours, no later than the end of the same business day.
- **impact/tier_2**: Important pipelines supporting established, actively used workflows that are not time-sensitive, such as external data feeds (app store reviews, Reddit, etc.), where a short delay is acceptable but an extended outage would affect planned work. A bug ticket must be created and the issue is expected to be resolved within 2–3 business days.
- **impact/tier_3**: No impact on other processes and is not used to generate any metrics used by business users or to make any decisions. A bug ticket should be created and it’s up to the job owner to fix this issue in whatever time frame they deem to be reasonable.

### triage/ tag
Expand Down
Loading