agora inbox for [email protected]help / color / mirror / Atom feed
[PATCH v26 7/9] Row pattern recognition patch (tests). 436+ messages / 2 participants [nested] [flat]
* [PATCH v26 7/9] Row pattern recognition patch (tests). @ 2024-12-30 12:44 Tatsuo Ishii <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Tatsuo Ishii @ 2024-12-30 12:44 UTC (permalink / raw) --- src/test/regress/expected/rpr.out | 919 +++++++++++++++++++++++++++++ src/test/regress/parallel_schedule | 2 +- src/test/regress/sql/rpr.sql | 467 +++++++++++++++ 3 files changed, 1387 insertions(+), 1 deletion(-) create mode 100644 src/test/regress/expected/rpr.out create mode 100644 src/test/regress/sql/rpr.sql diff --git a/src/test/regress/expected/rpr.out b/src/test/regress/expected/rpr.out new file mode 100644 index 0000000000..b8cd6190b4 --- /dev/null +++ b/src/test/regress/expected/rpr.out @@ -0,0 +1,919 @@ +-- +-- Test for row pattern definition clause +-- +CREATE TEMP TABLE stock ( + company TEXT, + tdate DATE, + price INTEGER +); +INSERT INTO stock VALUES ('company1', '2023-07-01', 100); +INSERT INTO stock VALUES ('company1', '2023-07-02', 200); +INSERT INTO stock VALUES ('company1', '2023-07-03', 150); +INSERT INTO stock VALUES ('company1', '2023-07-04', 140); +INSERT INTO stock VALUES ('company1', '2023-07-05', 150); +INSERT INTO stock VALUES ('company1', '2023-07-06', 90); +INSERT INTO stock VALUES ('company1', '2023-07-07', 110); +INSERT INTO stock VALUES ('company1', '2023-07-08', 130); +INSERT INTO stock VALUES ('company1', '2023-07-09', 120); +INSERT INTO stock VALUES ('company1', '2023-07-10', 130); +INSERT INTO stock VALUES ('company2', '2023-07-01', 50); +INSERT INTO stock VALUES ('company2', '2023-07-02', 2000); +INSERT INTO stock VALUES ('company2', '2023-07-03', 1500); +INSERT INTO stock VALUES ('company2', '2023-07-04', 1400); +INSERT INTO stock VALUES ('company2', '2023-07-05', 1500); +INSERT INTO stock VALUES ('company2', '2023-07-06', 60); +INSERT INTO stock VALUES ('company2', '2023-07-07', 1100); +INSERT INTO stock VALUES ('company2', '2023-07-08', 1300); +INSERT INTO stock VALUES ('company2', '2023-07-09', 1200); +INSERT INTO stock VALUES ('company2', '2023-07-10', 1300); +SELECT * FROM stock; + company | tdate | price +----------+------------+------- + company1 | 07-01-2023 | 100 + company1 | 07-02-2023 | 200 + company1 | 07-03-2023 | 150 + company1 | 07-04-2023 | 140 + company1 | 07-05-2023 | 150 + company1 | 07-06-2023 | 90 + company1 | 07-07-2023 | 110 + company1 | 07-08-2023 | 130 + company1 | 07-09-2023 | 120 + company1 | 07-10-2023 | 130 + company2 | 07-01-2023 | 50 + company2 | 07-02-2023 | 2000 + company2 | 07-03-2023 | 1500 + company2 | 07-04-2023 | 1400 + company2 | 07-05-2023 | 1500 + company2 | 07-06-2023 | 60 + company2 | 07-07-2023 | 1100 + company2 | 07-08-2023 | 1300 + company2 | 07-09-2023 | 1200 + company2 | 07-10-2023 | 1300 +(20 rows) + +-- basic test using PREV +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w, + nth_value(tdate, 2) OVER w AS nth_second + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + company | tdate | price | first_value | last_value | nth_second +----------+------------+-------+-------------+------------+------------ + company1 | 07-01-2023 | 100 | 100 | 140 | 07-02-2023 + company1 | 07-02-2023 | 200 | | | + company1 | 07-03-2023 | 150 | | | + company1 | 07-04-2023 | 140 | | | + company1 | 07-05-2023 | 150 | | | + company1 | 07-06-2023 | 90 | 90 | 120 | 07-07-2023 + company1 | 07-07-2023 | 110 | | | + company1 | 07-08-2023 | 130 | | | + company1 | 07-09-2023 | 120 | | | + company1 | 07-10-2023 | 130 | | | + company2 | 07-01-2023 | 50 | 50 | 1400 | 07-02-2023 + company2 | 07-02-2023 | 2000 | | | + company2 | 07-03-2023 | 1500 | | | + company2 | 07-04-2023 | 1400 | | | + company2 | 07-05-2023 | 1500 | | | + company2 | 07-06-2023 | 60 | 60 | 1200 | 07-07-2023 + company2 | 07-07-2023 | 1100 | | | + company2 | 07-08-2023 | 1300 | | | + company2 | 07-09-2023 | 1200 | | | + company2 | 07-10-2023 | 1300 | | | +(20 rows) + +-- basic test using PREV. UP appears twice +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w, + nth_value(tdate, 2) OVER w AS nth_second + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+ UP+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + company | tdate | price | first_value | last_value | nth_second +----------+------------+-------+-------------+------------+------------ + company1 | 07-01-2023 | 100 | 100 | 150 | 07-02-2023 + company1 | 07-02-2023 | 200 | | | + company1 | 07-03-2023 | 150 | | | + company1 | 07-04-2023 | 140 | | | + company1 | 07-05-2023 | 150 | | | + company1 | 07-06-2023 | 90 | 90 | 130 | 07-07-2023 + company1 | 07-07-2023 | 110 | | | + company1 | 07-08-2023 | 130 | | | + company1 | 07-09-2023 | 120 | | | + company1 | 07-10-2023 | 130 | | | + company2 | 07-01-2023 | 50 | 50 | 1500 | 07-02-2023 + company2 | 07-02-2023 | 2000 | | | + company2 | 07-03-2023 | 1500 | | | + company2 | 07-04-2023 | 1400 | | | + company2 | 07-05-2023 | 1500 | | | + company2 | 07-06-2023 | 60 | 60 | 1300 | 07-07-2023 + company2 | 07-07-2023 | 1100 | | | + company2 | 07-08-2023 | 1300 | | | + company2 | 07-09-2023 | 1200 | | | + company2 | 07-10-2023 | 1300 | | | +(20 rows) + +-- basic test using PREV. Use '*' +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w, + nth_value(tdate, 2) OVER w AS nth_second + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP* DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + company | tdate | price | first_value | last_value | nth_second +----------+------------+-------+-------------+------------+------------ + company1 | 07-01-2023 | 100 | 100 | 140 | 07-02-2023 + company1 | 07-02-2023 | 200 | | | + company1 | 07-03-2023 | 150 | | | + company1 | 07-04-2023 | 140 | | | + company1 | 07-05-2023 | 150 | 150 | 90 | 07-06-2023 + company1 | 07-06-2023 | 90 | | | + company1 | 07-07-2023 | 110 | 110 | 120 | 07-08-2023 + company1 | 07-08-2023 | 130 | | | + company1 | 07-09-2023 | 120 | | | + company1 | 07-10-2023 | 130 | | | + company2 | 07-01-2023 | 50 | 50 | 1400 | 07-02-2023 + company2 | 07-02-2023 | 2000 | | | + company2 | 07-03-2023 | 1500 | | | + company2 | 07-04-2023 | 1400 | | | + company2 | 07-05-2023 | 1500 | 1500 | 60 | 07-06-2023 + company2 | 07-06-2023 | 60 | | | + company2 | 07-07-2023 | 1100 | 1100 | 1200 | 07-08-2023 + company2 | 07-08-2023 | 1300 | | | + company2 | 07-09-2023 | 1200 | | | + company2 | 07-10-2023 | 1300 | | | +(20 rows) + +-- basic test with none greedy pattern +SELECT company, tdate, price, count(*) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (A A A) + DEFINE + A AS price >= 140 AND price <= 150 +); + company | tdate | price | count +----------+------------+-------+------- + company1 | 07-01-2023 | 100 | 0 + company1 | 07-02-2023 | 200 | 0 + company1 | 07-03-2023 | 150 | 3 + company1 | 07-04-2023 | 140 | 0 + company1 | 07-05-2023 | 150 | 0 + company1 | 07-06-2023 | 90 | 0 + company1 | 07-07-2023 | 110 | 0 + company1 | 07-08-2023 | 130 | 0 + company1 | 07-09-2023 | 120 | 0 + company1 | 07-10-2023 | 130 | 0 + company2 | 07-01-2023 | 50 | 0 + company2 | 07-02-2023 | 2000 | 0 + company2 | 07-03-2023 | 1500 | 0 + company2 | 07-04-2023 | 1400 | 0 + company2 | 07-05-2023 | 1500 | 0 + company2 | 07-06-2023 | 60 | 0 + company2 | 07-07-2023 | 1100 | 0 + company2 | 07-08-2023 | 1300 | 0 + company2 | 07-09-2023 | 1200 | 0 + company2 | 07-10-2023 | 1300 | 0 +(20 rows) + +-- last_value() should remain consistent +SELECT company, tdate, price, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + company | tdate | price | last_value +----------+------------+-------+------------ + company1 | 07-01-2023 | 100 | 140 + company1 | 07-02-2023 | 200 | + company1 | 07-03-2023 | 150 | + company1 | 07-04-2023 | 140 | + company1 | 07-05-2023 | 150 | + company1 | 07-06-2023 | 90 | 120 + company1 | 07-07-2023 | 110 | + company1 | 07-08-2023 | 130 | + company1 | 07-09-2023 | 120 | + company1 | 07-10-2023 | 130 | + company2 | 07-01-2023 | 50 | 1400 + company2 | 07-02-2023 | 2000 | + company2 | 07-03-2023 | 1500 | + company2 | 07-04-2023 | 1400 | + company2 | 07-05-2023 | 1500 | + company2 | 07-06-2023 | 60 | 1200 + company2 | 07-07-2023 | 1100 | + company2 | 07-08-2023 | 1300 | + company2 | 07-09-2023 | 1200 | + company2 | 07-10-2023 | 1300 | +(20 rows) + +-- omit "START" in DEFINE but it is ok because "START AS TRUE" is +-- implicitly defined. per spec. +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w, + nth_value(tdate, 2) OVER w AS nth_second + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + company | tdate | price | first_value | last_value | nth_second +----------+------------+-------+-------------+------------+------------ + company1 | 07-01-2023 | 100 | 100 | 140 | 07-02-2023 + company1 | 07-02-2023 | 200 | | | + company1 | 07-03-2023 | 150 | | | + company1 | 07-04-2023 | 140 | | | + company1 | 07-05-2023 | 150 | | | + company1 | 07-06-2023 | 90 | 90 | 120 | 07-07-2023 + company1 | 07-07-2023 | 110 | | | + company1 | 07-08-2023 | 130 | | | + company1 | 07-09-2023 | 120 | | | + company1 | 07-10-2023 | 130 | | | + company2 | 07-01-2023 | 50 | 50 | 1400 | 07-02-2023 + company2 | 07-02-2023 | 2000 | | | + company2 | 07-03-2023 | 1500 | | | + company2 | 07-04-2023 | 1400 | | | + company2 | 07-05-2023 | 1500 | | | + company2 | 07-06-2023 | 60 | 60 | 1200 | 07-07-2023 + company2 | 07-07-2023 | 1100 | | | + company2 | 07-08-2023 | 1300 | | | + company2 | 07-09-2023 | 1200 | | | + company2 | 07-10-2023 | 1300 | | | +(20 rows) + +-- the first row start with less than or equal to 100 +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (LOWPRICE UP+ DOWN+) + DEFINE + LOWPRICE AS price <= 100, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + company | tdate | price | first_value | last_value +----------+------------+-------+-------------+------------ + company1 | 07-01-2023 | 100 | 100 | 140 + company1 | 07-02-2023 | 200 | | + company1 | 07-03-2023 | 150 | | + company1 | 07-04-2023 | 140 | | + company1 | 07-05-2023 | 150 | | + company1 | 07-06-2023 | 90 | 90 | 120 + company1 | 07-07-2023 | 110 | | + company1 | 07-08-2023 | 130 | | + company1 | 07-09-2023 | 120 | | + company1 | 07-10-2023 | 130 | | + company2 | 07-01-2023 | 50 | 50 | 1400 + company2 | 07-02-2023 | 2000 | | + company2 | 07-03-2023 | 1500 | | + company2 | 07-04-2023 | 1400 | | + company2 | 07-05-2023 | 1500 | | + company2 | 07-06-2023 | 60 | 60 | 1200 + company2 | 07-07-2023 | 1100 | | + company2 | 07-08-2023 | 1300 | | + company2 | 07-09-2023 | 1200 | | + company2 | 07-10-2023 | 1300 | | +(20 rows) + +-- second row raises 120% +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (LOWPRICE UP+ DOWN+) + DEFINE + LOWPRICE AS price <= 100, + UP AS price > PREV(price) * 1.2, + DOWN AS price < PREV(price) +); + company | tdate | price | first_value | last_value +----------+------------+-------+-------------+------------ + company1 | 07-01-2023 | 100 | 100 | 140 + company1 | 07-02-2023 | 200 | | + company1 | 07-03-2023 | 150 | | + company1 | 07-04-2023 | 140 | | + company1 | 07-05-2023 | 150 | | + company1 | 07-06-2023 | 90 | | + company1 | 07-07-2023 | 110 | | + company1 | 07-08-2023 | 130 | | + company1 | 07-09-2023 | 120 | | + company1 | 07-10-2023 | 130 | | + company2 | 07-01-2023 | 50 | 50 | 1400 + company2 | 07-02-2023 | 2000 | | + company2 | 07-03-2023 | 1500 | | + company2 | 07-04-2023 | 1400 | | + company2 | 07-05-2023 | 1500 | | + company2 | 07-06-2023 | 60 | | + company2 | 07-07-2023 | 1100 | | + company2 | 07-08-2023 | 1300 | | + company2 | 07-09-2023 | 1200 | | + company2 | 07-10-2023 | 1300 | | +(20 rows) + +-- using NEXT +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UPDOWN) + DEFINE + START AS TRUE, + UPDOWN AS price > PREV(price) AND price > NEXT(price) +); + company | tdate | price | first_value | last_value +----------+------------+-------+-------------+------------ + company1 | 07-01-2023 | 100 | 100 | 200 + company1 | 07-02-2023 | 200 | | + company1 | 07-03-2023 | 150 | | + company1 | 07-04-2023 | 140 | 140 | 150 + company1 | 07-05-2023 | 150 | | + company1 | 07-06-2023 | 90 | | + company1 | 07-07-2023 | 110 | 110 | 130 + company1 | 07-08-2023 | 130 | | + company1 | 07-09-2023 | 120 | | + company1 | 07-10-2023 | 130 | | + company2 | 07-01-2023 | 50 | 50 | 2000 + company2 | 07-02-2023 | 2000 | | + company2 | 07-03-2023 | 1500 | | + company2 | 07-04-2023 | 1400 | 1400 | 1500 + company2 | 07-05-2023 | 1500 | | + company2 | 07-06-2023 | 60 | | + company2 | 07-07-2023 | 1100 | 1100 | 1300 + company2 | 07-08-2023 | 1300 | | + company2 | 07-09-2023 | 1200 | | + company2 | 07-10-2023 | 1300 | | +(20 rows) + +-- using AFTER MATCH SKIP TO NEXT ROW +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP TO NEXT ROW + INITIAL + PATTERN (START UPDOWN) + DEFINE + START AS TRUE, + UPDOWN AS price > PREV(price) AND price > NEXT(price) +); + company | tdate | price | first_value | last_value +----------+------------+-------+-------------+------------ + company1 | 07-01-2023 | 100 | 100 | 200 + company1 | 07-02-2023 | 200 | | + company1 | 07-03-2023 | 150 | | + company1 | 07-04-2023 | 140 | 140 | 150 + company1 | 07-05-2023 | 150 | | + company1 | 07-06-2023 | 90 | | + company1 | 07-07-2023 | 110 | 110 | 130 + company1 | 07-08-2023 | 130 | | + company1 | 07-09-2023 | 120 | | + company1 | 07-10-2023 | 130 | | + company2 | 07-01-2023 | 50 | 50 | 2000 + company2 | 07-02-2023 | 2000 | | + company2 | 07-03-2023 | 1500 | | + company2 | 07-04-2023 | 1400 | 1400 | 1500 + company2 | 07-05-2023 | 1500 | | + company2 | 07-06-2023 | 60 | | + company2 | 07-07-2023 | 1100 | 1100 | 1300 + company2 | 07-08-2023 | 1300 | | + company2 | 07-09-2023 | 1200 | | + company2 | 07-10-2023 | 1300 | | +(20 rows) + +-- match everything +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + INITIAL + PATTERN (A+) + DEFINE + A AS TRUE +); + company | tdate | price | first_value | last_value +----------+------------+-------+-------------+------------ + company1 | 07-01-2023 | 100 | 100 | 130 + company1 | 07-02-2023 | 200 | | + company1 | 07-03-2023 | 150 | | + company1 | 07-04-2023 | 140 | | + company1 | 07-05-2023 | 150 | | + company1 | 07-06-2023 | 90 | | + company1 | 07-07-2023 | 110 | | + company1 | 07-08-2023 | 130 | | + company1 | 07-09-2023 | 120 | | + company1 | 07-10-2023 | 130 | | + company2 | 07-01-2023 | 50 | 50 | 1300 + company2 | 07-02-2023 | 2000 | | + company2 | 07-03-2023 | 1500 | | + company2 | 07-04-2023 | 1400 | | + company2 | 07-05-2023 | 1500 | | + company2 | 07-06-2023 | 60 | | + company2 | 07-07-2023 | 1100 | | + company2 | 07-08-2023 | 1300 | | + company2 | 07-09-2023 | 1200 | | + company2 | 07-10-2023 | 1300 | | +(20 rows) + +-- backtracking with reclassification of rows +-- using AFTER MATCH SKIP PAST LAST ROW +SELECT company, tdate, price, first_value(tdate) OVER w, last_value(tdate) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + INITIAL + PATTERN (A+ B+) + DEFINE + A AS price > 100, + B AS price > 100 +); + company | tdate | price | first_value | last_value +----------+------------+-------+-------------+------------ + company1 | 07-01-2023 | 100 | | + company1 | 07-02-2023 | 200 | 07-02-2023 | 07-05-2023 + company1 | 07-03-2023 | 150 | | + company1 | 07-04-2023 | 140 | | + company1 | 07-05-2023 | 150 | | + company1 | 07-06-2023 | 90 | | + company1 | 07-07-2023 | 110 | 07-07-2023 | 07-10-2023 + company1 | 07-08-2023 | 130 | | + company1 | 07-09-2023 | 120 | | + company1 | 07-10-2023 | 130 | | + company2 | 07-01-2023 | 50 | | + company2 | 07-02-2023 | 2000 | 07-02-2023 | 07-05-2023 + company2 | 07-03-2023 | 1500 | | + company2 | 07-04-2023 | 1400 | | + company2 | 07-05-2023 | 1500 | | + company2 | 07-06-2023 | 60 | | + company2 | 07-07-2023 | 1100 | 07-07-2023 | 07-10-2023 + company2 | 07-08-2023 | 1300 | | + company2 | 07-09-2023 | 1200 | | + company2 | 07-10-2023 | 1300 | | +(20 rows) + +-- backtracking with reclassification of rows +-- using AFTER MATCH SKIP TO NEXT ROW +SELECT company, tdate, price, first_value(tdate) OVER w, last_value(tdate) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP TO NEXT ROW + INITIAL + PATTERN (A+ B+) + DEFINE + A AS price > 100, + B AS price > 100 +); + company | tdate | price | first_value | last_value +----------+------------+-------+-------------+------------ + company1 | 07-01-2023 | 100 | | + company1 | 07-02-2023 | 200 | 07-02-2023 | 07-05-2023 + company1 | 07-03-2023 | 150 | 07-03-2023 | 07-05-2023 + company1 | 07-04-2023 | 140 | 07-04-2023 | 07-05-2023 + company1 | 07-05-2023 | 150 | | + company1 | 07-06-2023 | 90 | | + company1 | 07-07-2023 | 110 | 07-07-2023 | 07-10-2023 + company1 | 07-08-2023 | 130 | 07-08-2023 | 07-10-2023 + company1 | 07-09-2023 | 120 | 07-09-2023 | 07-10-2023 + company1 | 07-10-2023 | 130 | | + company2 | 07-01-2023 | 50 | | + company2 | 07-02-2023 | 2000 | 07-02-2023 | 07-05-2023 + company2 | 07-03-2023 | 1500 | 07-03-2023 | 07-05-2023 + company2 | 07-04-2023 | 1400 | 07-04-2023 | 07-05-2023 + company2 | 07-05-2023 | 1500 | | + company2 | 07-06-2023 | 60 | | + company2 | 07-07-2023 | 1100 | 07-07-2023 | 07-10-2023 + company2 | 07-08-2023 | 1300 | 07-08-2023 | 07-10-2023 + company2 | 07-09-2023 | 1200 | 07-09-2023 | 07-10-2023 + company2 | 07-10-2023 | 1300 | | +(20 rows) + +-- ROWS BETWEEN CURRENT ROW AND offset FOLLOWING +SELECT company, tdate, price, first_value(tdate) OVER w, last_value(tdate) OVER w, + count(*) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND 2 FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + company | tdate | price | first_value | last_value | count +----------+------------+-------+-------------+------------+------- + company1 | 07-01-2023 | 100 | 07-01-2023 | 07-03-2023 | 3 + company1 | 07-02-2023 | 200 | | | 0 + company1 | 07-03-2023 | 150 | | | 0 + company1 | 07-04-2023 | 140 | 07-04-2023 | 07-06-2023 | 3 + company1 | 07-05-2023 | 150 | | | 0 + company1 | 07-06-2023 | 90 | | | 0 + company1 | 07-07-2023 | 110 | 07-07-2023 | 07-09-2023 | 3 + company1 | 07-08-2023 | 130 | | | 0 + company1 | 07-09-2023 | 120 | | | 0 + company1 | 07-10-2023 | 130 | | | 0 + company2 | 07-01-2023 | 50 | 07-01-2023 | 07-03-2023 | 3 + company2 | 07-02-2023 | 2000 | | | 0 + company2 | 07-03-2023 | 1500 | | | 0 + company2 | 07-04-2023 | 1400 | 07-04-2023 | 07-06-2023 | 3 + company2 | 07-05-2023 | 1500 | | | 0 + company2 | 07-06-2023 | 60 | | | 0 + company2 | 07-07-2023 | 1100 | 07-07-2023 | 07-09-2023 | 3 + company2 | 07-08-2023 | 1300 | | | 0 + company2 | 07-09-2023 | 1200 | | | 0 + company2 | 07-10-2023 | 1300 | | | 0 +(20 rows) + +-- +-- Aggregates +-- +-- using AFTER MATCH SKIP PAST LAST ROW +SELECT company, tdate, price, + first_value(price) OVER w, + last_value(price) OVER w, + max(price) OVER w, + min(price) OVER w, + sum(price) OVER w, + avg(price) OVER w, + count(price) OVER w +FROM stock +WINDOW w AS ( +PARTITION BY company +ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING +AFTER MATCH SKIP PAST LAST ROW +INITIAL +PATTERN (START UP+ DOWN+) +DEFINE +START AS TRUE, +UP AS price > PREV(price), +DOWN AS price < PREV(price) +); + company | tdate | price | first_value | last_value | max | min | sum | avg | count +----------+------------+-------+-------------+------------+------+-----+------+-----------------------+------- + company1 | 07-01-2023 | 100 | 100 | 140 | 200 | 100 | 590 | 147.5000000000000000 | 4 + company1 | 07-02-2023 | 200 | | | | | | | 0 + company1 | 07-03-2023 | 150 | | | | | | | 0 + company1 | 07-04-2023 | 140 | | | | | | | 0 + company1 | 07-05-2023 | 150 | | | | | | | 0 + company1 | 07-06-2023 | 90 | 90 | 120 | 130 | 90 | 450 | 112.5000000000000000 | 4 + company1 | 07-07-2023 | 110 | | | | | | | 0 + company1 | 07-08-2023 | 130 | | | | | | | 0 + company1 | 07-09-2023 | 120 | | | | | | | 0 + company1 | 07-10-2023 | 130 | | | | | | | 0 + company2 | 07-01-2023 | 50 | 50 | 1400 | 2000 | 50 | 4950 | 1237.5000000000000000 | 4 + company2 | 07-02-2023 | 2000 | | | | | | | 0 + company2 | 07-03-2023 | 1500 | | | | | | | 0 + company2 | 07-04-2023 | 1400 | | | | | | | 0 + company2 | 07-05-2023 | 1500 | | | | | | | 0 + company2 | 07-06-2023 | 60 | 60 | 1200 | 1300 | 60 | 3660 | 915.0000000000000000 | 4 + company2 | 07-07-2023 | 1100 | | | | | | | 0 + company2 | 07-08-2023 | 1300 | | | | | | | 0 + company2 | 07-09-2023 | 1200 | | | | | | | 0 + company2 | 07-10-2023 | 1300 | | | | | | | 0 +(20 rows) + +-- using AFTER MATCH SKIP TO NEXT ROW +SELECT company, tdate, price, + first_value(price) OVER w, + last_value(price) OVER w, + max(price) OVER w, + min(price) OVER w, + sum(price) OVER w, + avg(price) OVER w, + count(price) OVER w +FROM stock +WINDOW w AS ( +PARTITION BY company +ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING +AFTER MATCH SKIP TO NEXT ROW +INITIAL +PATTERN (START UP+ DOWN+) +DEFINE +START AS TRUE, +UP AS price > PREV(price), +DOWN AS price < PREV(price) +); + company | tdate | price | first_value | last_value | max | min | sum | avg | count +----------+------------+-------+-------------+------------+------+------+------+-----------------------+------- + company1 | 07-01-2023 | 100 | 100 | 140 | 200 | 100 | 590 | 147.5000000000000000 | 4 + company1 | 07-02-2023 | 200 | | | | | | | 0 + company1 | 07-03-2023 | 150 | | | | | | | 0 + company1 | 07-04-2023 | 140 | 140 | 90 | 150 | 90 | 380 | 126.6666666666666667 | 3 + company1 | 07-05-2023 | 150 | | | | | | | 0 + company1 | 07-06-2023 | 90 | 90 | 120 | 130 | 90 | 450 | 112.5000000000000000 | 4 + company1 | 07-07-2023 | 110 | 110 | 120 | 130 | 110 | 360 | 120.0000000000000000 | 3 + company1 | 07-08-2023 | 130 | | | | | | | 0 + company1 | 07-09-2023 | 120 | | | | | | | 0 + company1 | 07-10-2023 | 130 | | | | | | | 0 + company2 | 07-01-2023 | 50 | 50 | 1400 | 2000 | 50 | 4950 | 1237.5000000000000000 | 4 + company2 | 07-02-2023 | 2000 | | | | | | | 0 + company2 | 07-03-2023 | 1500 | | | | | | | 0 + company2 | 07-04-2023 | 1400 | 1400 | 60 | 1500 | 60 | 2960 | 986.6666666666666667 | 3 + company2 | 07-05-2023 | 1500 | | | | | | | 0 + company2 | 07-06-2023 | 60 | 60 | 1200 | 1300 | 60 | 3660 | 915.0000000000000000 | 4 + company2 | 07-07-2023 | 1100 | 1100 | 1200 | 1300 | 1100 | 3600 | 1200.0000000000000000 | 3 + company2 | 07-08-2023 | 1300 | | | | | | | 0 + company2 | 07-09-2023 | 1200 | | | | | | | 0 + company2 | 07-10-2023 | 1300 | | | | | | | 0 +(20 rows) + +-- JOIN case +CREATE TEMP TABLE t1 (i int, v1 int); +CREATE TEMP TABLE t2 (j int, v2 int); +INSERT INTO t1 VALUES(1,10); +INSERT INTO t1 VALUES(1,11); +INSERT INTO t1 VALUES(1,12); +INSERT INTO t2 VALUES(2,10); +INSERT INTO t2 VALUES(2,11); +INSERT INTO t2 VALUES(2,12); +SELECT * FROM t1, t2 WHERE t1.v1 <= 11 AND t2.v2 <= 11; + i | v1 | j | v2 +---+----+---+---- + 1 | 10 | 2 | 10 + 1 | 10 | 2 | 11 + 1 | 11 | 2 | 10 + 1 | 11 | 2 | 11 +(4 rows) + +SELECT *, count(*) OVER w FROM t1, t2 +WINDOW w AS ( + PARTITION BY t1.i + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (A) + DEFINE + A AS v1 <= 11 AND v2 <= 11 +); + i | v1 | j | v2 | count +---+----+---+----+------- + 1 | 10 | 2 | 10 | 1 + 1 | 10 | 2 | 11 | 1 + 1 | 10 | 2 | 12 | 0 + 1 | 11 | 2 | 10 | 1 + 1 | 11 | 2 | 11 | 1 + 1 | 11 | 2 | 12 | 0 + 1 | 12 | 2 | 10 | 0 + 1 | 12 | 2 | 11 | 0 + 1 | 12 | 2 | 12 | 0 +(9 rows) + +-- WITH case +WITH wstock AS ( + SELECT * FROM stock WHERE tdate < '2023-07-08' +) +SELECT tdate, price, +first_value(tdate) OVER w, +count(*) OVER w + FROM wstock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + tdate | price | first_value | count +------------+-------+-------------+------- + 07-01-2023 | 100 | 07-01-2023 | 4 + 07-02-2023 | 200 | | 0 + 07-03-2023 | 150 | | 0 + 07-04-2023 | 140 | | 0 + 07-05-2023 | 150 | | 0 + 07-06-2023 | 90 | | 0 + 07-07-2023 | 110 | | 0 + 07-01-2023 | 50 | 07-01-2023 | 4 + 07-02-2023 | 2000 | | 0 + 07-03-2023 | 1500 | | 0 + 07-04-2023 | 1400 | | 0 + 07-05-2023 | 1500 | | 0 + 07-06-2023 | 60 | | 0 + 07-07-2023 | 1100 | | 0 +(14 rows) + +-- PREV has multiple column reference +CREATE TEMP TABLE rpr1 (id INTEGER, i SERIAL, j INTEGER); +INSERT INTO rpr1(id, j) SELECT 1, g*2 FROM generate_series(1, 10) AS g; +SELECT id, i, j, count(*) OVER w + FROM rpr1 + WINDOW w AS ( + PARTITION BY id + ORDER BY i + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + INITIAL + PATTERN (START COND+) + DEFINE + START AS TRUE, + COND AS PREV(i + j + 1) < 10 +); + id | i | j | count +----+----+----+------- + 1 | 1 | 2 | 3 + 1 | 2 | 4 | 0 + 1 | 3 | 6 | 0 + 1 | 4 | 8 | 0 + 1 | 5 | 10 | 0 + 1 | 6 | 12 | 0 + 1 | 7 | 14 | 0 + 1 | 8 | 16 | 0 + 1 | 9 | 18 | 0 + 1 | 10 | 20 | 0 +(10 rows) + +-- Smoke test for larger partitions. +WITH s AS ( + SELECT v, count(*) OVER w AS c + FROM (SELECT generate_series(1, 5000) v) + WINDOW w AS ( + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + INITIAL + PATTERN ( r+ ) + DEFINE r AS TRUE + ) +) +-- Should be exactly one long match across all rows. +SELECT * FROM s WHERE c > 0; + v | c +---+------ + 1 | 5000 +(1 row) + +WITH s AS ( + SELECT v, count(*) OVER w AS c + FROM (SELECT generate_series(1, 5000) v) + WINDOW w AS ( + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + INITIAL + PATTERN ( r ) + DEFINE r AS TRUE + ) +) +-- Every row should be its own match. +SELECT count(*) FROM s WHERE c > 0; + count +------- + 5000 +(1 row) + +-- +-- Error cases +-- +-- row pattern definition variable name must not appear more than once +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price), + UP AS price > PREV(price) +); +ERROR: row pattern definition variable name "up" appears more than once in DEFINE clause +LINE 11: UP AS price > PREV(price), + ^ +-- subqueries in DEFINE clause are not supported +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START LOWPRICE) + DEFINE + START AS TRUE, + LOWPRICE AS price < (SELECT 100) +); +ERROR: cannot use subquery in DEFINE expression +LINE 11: LOWPRICE AS price < (SELECT 100) + ^ +-- aggregates in DEFINE clause are not supported +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START LOWPRICE) + DEFINE + START AS TRUE, + LOWPRICE AS price < count(*) +); +ERROR: aggregate functions are not allowed in DEFINE +LINE 11: LOWPRICE AS price < count(*) + ^ +-- FRAME must start at current row when row patttern recognition is used +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN UNBOUNDED PRECEDING AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); +ERROR: FRAME must start at current row when row patttern recognition is used +-- SEEK is not supported +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP TO NEXT ROW + SEEK + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); +ERROR: SEEK is not supported +LINE 8: SEEK + ^ +HINT: Use INITIAL. +-- PREV's argument must have at least 1 column reference +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP TO NEXT ROW + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(1), + DOWN AS price < PREV(1) +); +ERROR: row pattern navigation operation's argument must include at least one column reference diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule index 1edd9e45eb..f5d33b4d7d 100644 --- a/src/test/regress/parallel_schedule +++ b/src/test/regress/parallel_schedule @@ -98,7 +98,7 @@ test: publication subscription # Another group of parallel tests # select_views depends on create_view # ---------- -test: select_views portals_p2 foreign_key cluster dependency guc bitmapops combocid tsearch tsdicts foreign_data window xmlmap functional_deps advisory_lock indirect_toast equivclass +test: select_views portals_p2 foreign_key cluster dependency guc bitmapops combocid tsearch tsdicts foreign_data window xmlmap functional_deps advisory_lock indirect_toast equivclass rpr # ---------- # Another group of parallel tests (JSON related) diff --git a/src/test/regress/sql/rpr.sql b/src/test/regress/sql/rpr.sql new file mode 100644 index 0000000000..a46abe6f0f --- /dev/null +++ b/src/test/regress/sql/rpr.sql @@ -0,0 +1,467 @@ +-- +-- Test for row pattern definition clause +-- + +CREATE TEMP TABLE stock ( + company TEXT, + tdate DATE, + price INTEGER +); +INSERT INTO stock VALUES ('company1', '2023-07-01', 100); +INSERT INTO stock VALUES ('company1', '2023-07-02', 200); +INSERT INTO stock VALUES ('company1', '2023-07-03', 150); +INSERT INTO stock VALUES ('company1', '2023-07-04', 140); +INSERT INTO stock VALUES ('company1', '2023-07-05', 150); +INSERT INTO stock VALUES ('company1', '2023-07-06', 90); +INSERT INTO stock VALUES ('company1', '2023-07-07', 110); +INSERT INTO stock VALUES ('company1', '2023-07-08', 130); +INSERT INTO stock VALUES ('company1', '2023-07-09', 120); +INSERT INTO stock VALUES ('company1', '2023-07-10', 130); +INSERT INTO stock VALUES ('company2', '2023-07-01', 50); +INSERT INTO stock VALUES ('company2', '2023-07-02', 2000); +INSERT INTO stock VALUES ('company2', '2023-07-03', 1500); +INSERT INTO stock VALUES ('company2', '2023-07-04', 1400); +INSERT INTO stock VALUES ('company2', '2023-07-05', 1500); +INSERT INTO stock VALUES ('company2', '2023-07-06', 60); +INSERT INTO stock VALUES ('company2', '2023-07-07', 1100); +INSERT INTO stock VALUES ('company2', '2023-07-08', 1300); +INSERT INTO stock VALUES ('company2', '2023-07-09', 1200); +INSERT INTO stock VALUES ('company2', '2023-07-10', 1300); + +SELECT * FROM stock; + +-- basic test using PREV +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w, + nth_value(tdate, 2) OVER w AS nth_second + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + +-- basic test using PREV. UP appears twice +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w, + nth_value(tdate, 2) OVER w AS nth_second + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+ UP+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + +-- basic test using PREV. Use '*' +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w, + nth_value(tdate, 2) OVER w AS nth_second + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP* DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + +-- basic test with none greedy pattern +SELECT company, tdate, price, count(*) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (A A A) + DEFINE + A AS price >= 140 AND price <= 150 +); + +-- last_value() should remain consistent +SELECT company, tdate, price, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + +-- omit "START" in DEFINE but it is ok because "START AS TRUE" is +-- implicitly defined. per spec. +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w, + nth_value(tdate, 2) OVER w AS nth_second + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + +-- the first row start with less than or equal to 100 +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (LOWPRICE UP+ DOWN+) + DEFINE + LOWPRICE AS price <= 100, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + +-- second row raises 120% +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (LOWPRICE UP+ DOWN+) + DEFINE + LOWPRICE AS price <= 100, + UP AS price > PREV(price) * 1.2, + DOWN AS price < PREV(price) +); + +-- using NEXT +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UPDOWN) + DEFINE + START AS TRUE, + UPDOWN AS price > PREV(price) AND price > NEXT(price) +); + +-- using AFTER MATCH SKIP TO NEXT ROW +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP TO NEXT ROW + INITIAL + PATTERN (START UPDOWN) + DEFINE + START AS TRUE, + UPDOWN AS price > PREV(price) AND price > NEXT(price) +); + +-- match everything + +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + INITIAL + PATTERN (A+) + DEFINE + A AS TRUE +); + +-- backtracking with reclassification of rows +-- using AFTER MATCH SKIP PAST LAST ROW +SELECT company, tdate, price, first_value(tdate) OVER w, last_value(tdate) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + INITIAL + PATTERN (A+ B+) + DEFINE + A AS price > 100, + B AS price > 100 +); + +-- backtracking with reclassification of rows +-- using AFTER MATCH SKIP TO NEXT ROW +SELECT company, tdate, price, first_value(tdate) OVER w, last_value(tdate) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP TO NEXT ROW + INITIAL + PATTERN (A+ B+) + DEFINE + A AS price > 100, + B AS price > 100 +); + +-- ROWS BETWEEN CURRENT ROW AND offset FOLLOWING +SELECT company, tdate, price, first_value(tdate) OVER w, last_value(tdate) OVER w, + count(*) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND 2 FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + +-- +-- Aggregates +-- + +-- using AFTER MATCH SKIP PAST LAST ROW +SELECT company, tdate, price, + first_value(price) OVER w, + last_value(price) OVER w, + max(price) OVER w, + min(price) OVER w, + sum(price) OVER w, + avg(price) OVER w, + count(price) OVER w +FROM stock +WINDOW w AS ( +PARTITION BY company +ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING +AFTER MATCH SKIP PAST LAST ROW +INITIAL +PATTERN (START UP+ DOWN+) +DEFINE +START AS TRUE, +UP AS price > PREV(price), +DOWN AS price < PREV(price) +); + +-- using AFTER MATCH SKIP TO NEXT ROW +SELECT company, tdate, price, + first_value(price) OVER w, + last_value(price) OVER w, + max(price) OVER w, + min(price) OVER w, + sum(price) OVER w, + avg(price) OVER w, + count(price) OVER w +FROM stock +WINDOW w AS ( +PARTITION BY company +ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING +AFTER MATCH SKIP TO NEXT ROW +INITIAL +PATTERN (START UP+ DOWN+) +DEFINE +START AS TRUE, +UP AS price > PREV(price), +DOWN AS price < PREV(price) +); + +-- JOIN case +CREATE TEMP TABLE t1 (i int, v1 int); +CREATE TEMP TABLE t2 (j int, v2 int); +INSERT INTO t1 VALUES(1,10); +INSERT INTO t1 VALUES(1,11); +INSERT INTO t1 VALUES(1,12); +INSERT INTO t2 VALUES(2,10); +INSERT INTO t2 VALUES(2,11); +INSERT INTO t2 VALUES(2,12); + +SELECT * FROM t1, t2 WHERE t1.v1 <= 11 AND t2.v2 <= 11; + +SELECT *, count(*) OVER w FROM t1, t2 +WINDOW w AS ( + PARTITION BY t1.i + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (A) + DEFINE + A AS v1 <= 11 AND v2 <= 11 +); + +-- WITH case +WITH wstock AS ( + SELECT * FROM stock WHERE tdate < '2023-07-08' +) +SELECT tdate, price, +first_value(tdate) OVER w, +count(*) OVER w + FROM wstock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + +-- PREV has multiple column reference +CREATE TEMP TABLE rpr1 (id INTEGER, i SERIAL, j INTEGER); +INSERT INTO rpr1(id, j) SELECT 1, g*2 FROM generate_series(1, 10) AS g; +SELECT id, i, j, count(*) OVER w + FROM rpr1 + WINDOW w AS ( + PARTITION BY id + ORDER BY i + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + INITIAL + PATTERN (START COND+) + DEFINE + START AS TRUE, + COND AS PREV(i + j + 1) < 10 +); + +-- Smoke test for larger partitions. +WITH s AS ( + SELECT v, count(*) OVER w AS c + FROM (SELECT generate_series(1, 5000) v) + WINDOW w AS ( + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + INITIAL + PATTERN ( r+ ) + DEFINE r AS TRUE + ) +) +-- Should be exactly one long match across all rows. +SELECT * FROM s WHERE c > 0; + +WITH s AS ( + SELECT v, count(*) OVER w AS c + FROM (SELECT generate_series(1, 5000) v) + WINDOW w AS ( + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP PAST LAST ROW + INITIAL + PATTERN ( r ) + DEFINE r AS TRUE + ) +) +-- Every row should be its own match. +SELECT count(*) FROM s WHERE c > 0; + +-- +-- Error cases +-- + +-- row pattern definition variable name must not appear more than once +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price), + UP AS price > PREV(price) +); + +-- subqueries in DEFINE clause are not supported +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START LOWPRICE) + DEFINE + START AS TRUE, + LOWPRICE AS price < (SELECT 100) +); + +-- aggregates in DEFINE clause are not supported +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START LOWPRICE) + DEFINE + START AS TRUE, + LOWPRICE AS price < count(*) +); + +-- FRAME must start at current row when row patttern recognition is used +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN UNBOUNDED PRECEDING AND UNBOUNDED FOLLOWING + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + +-- SEEK is not supported +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP TO NEXT ROW + SEEK + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(price), + DOWN AS price < PREV(price) +); + +-- PREV's argument must have at least 1 column reference +SELECT company, tdate, price, first_value(price) OVER w, last_value(price) OVER w + FROM stock + WINDOW w AS ( + PARTITION BY company + ORDER BY tdate + ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING + AFTER MATCH SKIP TO NEXT ROW + INITIAL + PATTERN (START UP+ DOWN+) + DEFINE + START AS TRUE, + UP AS price > PREV(1), + DOWN AS price < PREV(1) +); -- 2.25.1 ----Next_Part(Mon_Dec_30_22_37_18_2024_171)-- Content-Type: Text/X-Patch; charset=us-ascii Content-Transfer-Encoding: 7bit Content-Disposition: inline; filename="v26-0008-Row-pattern-recognition-patch-typedefs.list.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
* [PATCH v3 1/2] ci: Improve ccache handling @ 2026-06-05 15:58 Andres Freund <[email protected]> 0 siblings, 0 replies; 436+ messages in thread From: Andres Freund @ 2026-06-05 15:58 UTC (permalink / raw) There previously were a number of issues: - We'd upload the cache even if we already had a high hit rate. That means we churn through the available cache space very quickly. For this we now check if the cache hit ratio is already high, and skip uploading a new cache in that case. - We'd generate per-branch caches, even if master's already would suffice, because the branch doesn't change much This is solved indirectly by the above. - The cache key allowed prefix matches based on the branch, e.g. master-pending would always use master's branch Replace the cache key element separator of - with :, which is not a valid part of a branch name. - When rebasing a feature branch, we'd start with just that branch's cache, rather than also having the newer cache of master available This is solved by downloading by master's and the feature branch's cache, simply overlaying both. That's possible because ccache is content addressed. - The size of a cache would increase to the max, even though there likely will be no benefit from old cache entries. Address this by explicitly evicting old data and also recompressing the cache before uploading it. In my testing this utilizes the available cache space (10GB for personal accounts) much more effictively than before. The not entirely trivial determination of whether it's worth uploading a cache entry is moved to a python script. I first had it as shell, but that gets awkward. This way it'd also be more viable to use ccache for msvc at some point. The per-job redundancies are a bit annoying. There's a way around that, by using composite actions, but I think that might be harder to understand, without all that much of an improvement. --- .github/workflows/pg-ci.yml | 94 ++++++++++++++++++++++------ src/tools/ci/gha_ccache_decide.py | 100 ++++++++++++++++++++++++++++++ 2 files changed, 176 insertions(+), 18 deletions(-) create mode 100644 src/tools/ci/gha_ccache_decide.py diff --git a/.github/workflows/pg-ci.yml b/.github/workflows/pg-ci.yml index 8560e9389f6..86dc47de8db 100644 --- a/.github/workflows/pg-ci.yml +++ b/.github/workflows/pg-ci.yml @@ -130,6 +130,22 @@ env: # commit-message directive parsed in the `setup` job below. CI_OS_ONLY_JOBS: "linux macos windows mingw compilerwarnings sanitycheck" + ### + # A few variables to make expressions later on shorter + ### + + ON_DEFAULT_BRANCH: ${{github.event.repository.default_branch == github.ref_name }} + + # Note that we need to be careful to use a separator that can't be in branch + # names, otherwise e.g. caches for 'master' might be restored on the + # 'master-pending' branch. + CACHE_PREFIX_DEFAULT: >- + :${{ github.job }}:${{ github.event.repository.default_branch }}: + CACHE_PREFIX_BRANCH: >- + :${{ github.job }}:${{ github.ref_name }}: + CACHE_SUFFIX: >- + ${{ github.run_id }}:${{ github.run_attempt }} + jobs: @@ -277,16 +293,30 @@ jobs: with: fetch-depth: ${{ env.CLONE_DEPTH }} - - &ccache_restore_step - name: Restore ccache - id: ccache_restore + # We restore both the ccache from the default branch (typically master), + # and from the current branch. This will often allow feature branches to + # start out with a high cache hit ratio. + # + # With ccache it turns out to work to just restore two caches into the + # same directory, as it's basically a content addressed store. Stats + # could be corrupted, but we zero them out anyway. + - &ccache_restore_default_step + name: "ccache: Restore for default branch ${{github.event.repository.default_branch}}" + if: ${{ env.ON_DEFAULT_BRANCH == 'false' }} uses: actions/cache/restore@v5 with: path: ${{ env.CCACHE_DIR }} - key: ccache-${{ github.job }}-${{ github.ref_name }}-${{ github.run_id }}-${{ github.run_attempt }} - restore-keys: | - ccache-${{ github.job }}-${{ github.ref_name }}- - ccache-${{ github.job }}- + key: ccache${{env.CACHE_PREFIX_DEFAULT}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_DEFAULT}} + + - &ccache_restore_branch_step + name: "ccache: Restore for branch ${{ github.ref_name }}" + id: ccache-restore-branch + uses: actions/cache/restore@v5 + with: + path: ${{ env.CCACHE_DIR }} + key: ccache${{env.CACHE_PREFIX_BRANCH}}${{env.CACHE_SUFFIX}} + restore-keys: ccache${{env.CACHE_PREFIX_BRANCH}} - &linux_prepare_workspace_step name: Prepare workspace @@ -325,15 +355,30 @@ jobs: ninja -C build -j${{env.BUILD_JOBS}} ${{env.MBUILD_TARGET}} ninja -C build -t missingdeps - # TODO: As long as we use per-run ccache caches, we should probably add - # a step that checks if there is sufficient new content to warrant - # saving the new cache. + # Decide if it's worth uploading a new version of the ccache cache. If + # we always do so unconditionally, we'd very quickly go through the + # allowed cache space. Instead we check if the hit rate is high enough + # already for that not to be worth it. + - &ccache_decide_save_step + name: "ccache: Decide if cache should be uploaded" + id: ccache-pre-save + # [Decide to] store the cache whenever the cache was set up, so that + # incrementally addressing compiler errors/warnings doesn't have to + # start from scratch. + if: | + always() && + steps.ccache-restore-branch.conclusion == 'success' + run: python3 src/tools/ci/gha_ccache_decide.py + - &ccache_save_step - name: Save ccache + name: "ccache: Upload cache" uses: actions/cache/save@v5 + if: | + always() && + steps.ccache-pre-save.outputs.should_save == 'true' with: path: ${{ env.CCACHE_DIR }} - key: ${{ steps.ccache_restore.outputs.cache-primary-key }} + key: ${{ steps.ccache-restore-branch.outputs.cache-primary-key }} # Run a minimal set of tests. The main regression tests take too long # for this purpose. For now this is a random quick pg_regress style @@ -448,7 +493,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -467,6 +513,7 @@ jobs: run: | make -s -j${BUILD_JOBS} world-bin + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -508,7 +555,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -527,6 +575,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -596,7 +645,8 @@ jobs: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - *linux_prepare_workspace_step - name: Configure @@ -613,6 +663,7 @@ jobs: shell: *su_postgres_shell run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -682,7 +733,6 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step - name: Setup core files run: | @@ -745,6 +795,9 @@ jobs: path: ${{ env.MACPORTS_CACHE }} key: ${{ steps.mp-key.outputs.key }} + - *ccache_restore_default_step + - *ccache_restore_branch_step + - name: Configure env: PKG_CONFIG_PATH: /opt/local/lib/pkgconfig/ @@ -762,6 +815,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1062,7 +1116,8 @@ jobs: shell: cmd run: mkdir ${{env.PG_REGRESS_SOCK_DIR}} - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Configure run: | @@ -1077,6 +1132,7 @@ jobs: - name: Build run: *ninja_build_cmd + - *ccache_decide_save_step - *ccache_save_step - name: Test world @@ -1118,7 +1174,8 @@ jobs: steps: - *nix_sysinfo_step - *checkout_step - - *ccache_restore_step + - *ccache_restore_default_step + - *ccache_restore_branch_step - name: Setup workspace run: | @@ -1213,5 +1270,6 @@ jobs: headerscheck cpluspluscheck \ EXTRAFLAGS='-fmax-errors=10' + - *ccache_decide_save_step - *ccache_save_step - *upload_logs_step diff --git a/src/tools/ci/gha_ccache_decide.py b/src/tools/ci/gha_ccache_decide.py new file mode 100644 index 00000000000..920f7bf9685 --- /dev/null +++ b/src/tools/ci/gha_ccache_decide.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 + +import os +import re +import shutil +import subprocess + +def run(cmd, check=True): + return subprocess.run( + cmd, + check=check, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + ).stdout + +def parse_ccache_stats(): + out = run(["ccache", "--print-stats"]) + hits = 0 + misses = 0 + + for line in out.splitlines(): + line = line.strip() + m = re.match(r"^local_storage_hit\s+(\d+)$", line) + if m: + hits = int(m.group(1)) + continue + m = re.match(r"^local_storage_miss\s+(\d+)$", line) + if m: + misses = int(m.group(1)) + continue + + return hits, misses + +def append_github_output(key, value): + output_path = os.environ["GITHUB_OUTPUT"] + with open(output_path, "a", encoding="utf-8") as f: + f.write(f"{key}={value}\n") + +def main(): + on_default_branch = os.environ["ON_DEFAULT_BRANCH"] == "true" + ccache_dir = os.environ["CCACHE_DIR"] + + # Decide the target hit percentage below which we decide to upload a new + # cache. On non-default branches a few misses aren't that bad. But, as the + # caches of the default branch are shared with all branches, it's worth + # aiming for a higher ratio there. + target_rate = 95 if on_default_branch else 80 + + # Log ccache stats, useful for more in-depth understanding. The avoid it + # swamping the output, collapse it in a group. + print("::group::ccache_stats") + print(run(["ccache", "-s", "-vv"])) + print("::endgroup::") + + # compute cache hit ratio + hits, misses = parse_ccache_stats() + total = hits + misses + hit_pct = int(( hits / total) * 100) if total > 0 else 100 + + print(f"hits: {hits}, misses: {misses}, hit_pct: {hit_pct}, target rate: {target_rate}") + + # If there were either barely any misses, or the cache hit ratio was high, + # there no point in generating a new cache entry. We have limited cache + # space. + should_save = misses > 10 and hit_pct < target_rate + + append_github_output("should_save", str(should_save).lower()) + + if not should_save: + print(f"hit rate {hit_pct} is above target of {target_rate}, skip creating new cache entry") + return 0 + + print(f"hit rate {hit_pct} is below target of {target_rate}, create new cache entry") + + # It's not worth persisting old cache entries (e.g. from before a + # change to a central header, or from the default branch if this + # branch differs a lot). Therefore evict ccache entries that are a + # bit older. The cutoff here is fairly arbitrary, it could + # probably be improved. + print("::group::ccache_shrink") + print(run(["ccache", "--evict-older-than", f"{45*60}s"])) + print(run(["ccache", "-X", "10"])) + + # Don't store ccache stats , otherwise we'd need to reset the cache access + # data after restoring the cache in the next run, to be able to get the + # hit ratio of the CI run. + print(run(["ccache", "-z"])) + print("::endgroup::") + + # Before continuing, try to kill all ccache instances, otherwise + # it's possible that on cancellations there is still running + # ccaches that cause the upload to fail. + if shutil.which("killall"): + print(run(["killall", "ccache"], check=False)) + + return 0 + +if __name__ == "__main__": + exit(main()) -- 2.54.0.380.gc69baaf57b --zzyr2ozda5zjkkci Content-Type: text/x-diff; charset=us-ascii Content-Disposition: attachment; filename="v3-0002-ci-fewer-tests.patch" ^ permalink raw reply [nested|flat] 436+ messages in thread
end of thread, other threads:[~2026-06-05 15:58 UTC | newest] Thread overview: 436+ messages (download: mbox mbox.gz follow: Atom feed) -- links below jump to the message on this page -- 2024-12-30 12:44 [PATCH v26 7/9] Row pattern recognition patch (tests). Tatsuo Ishii <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]> 2026-06-05 15:58 [PATCH v3 1/2] ci: Improve ccache handling Andres Freund <[email protected]>
This inbox is served by agora; see mirroring instructions for how to clone and mirror all data and code used for this inbox